Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities

July 26, 2026 Β· View on GitHub

This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities

[πŸ“– arXiv Paper] [πŸ’» Project Page] [πŸ“Š Dataset] [πŸ† Leaderboard]

Examples of Daily-Omni Benchmark

News

  • [2026-07-26] Added AGIBOT X-Lab WITA-Omni Preview, Qwen3.5-Omni-Plus, Gemini 3.1 Pro Preview, and Doubao Seed 2.0 Lite to the audio-visual leaderboard. We thank AGIBOT for providing the evaluation results for Gemini, Doubao, and Qwen, and congratulate AGIBOT X-Lab WITA-Omni Preview on achieving the best performance on Daily-Omni!
  • [2026-04-27] Added NVIDIA Nemotron 3 Nano Omni 30B A3B to the leaderboard as the current top evaluated open-weight omni-modal model.
  • [2026-03-11] Updated the arXiv paper and refreshed the Leaderboard with new results.

Leaderboard

Performance comparison of MLLMs on Daily-Omni. Random guess accuracy is 25%.

Abbreviations:

  • AV Align: audio-visual alignment
  • Comp.: comparative
  • Ctx. Und.: context understanding
  • Evt. Seq.: event sequence
  • Infer.: inference
  • Reas.: reasoning
  • 30s / 60s: duration-based subsets

Closed-source models are marked with (Closed) and open-source models with (Open).

Omni-Modal Language Models (With Visual and Audio)

MethodsAV AlignComparisonContext UnderstandingEvent SequenceInferenceReasoning30s60sAvg
AGIBOT X-Lab WITA-Omni Preview (Closed)86.1389.3181.8784.6485.0685.7183.6287.0985.21
Qwen3.5-Omni-Plus (Closed)†84.4288.5582.7283.6784.6786.0583.4486.2084.68
Gemini 3.1 Pro Preview (Closed)83.6186.2679.7982.3581.1784.5781.7684.0082.79
Doubao Seed 2.0 Lite (Closed)81.5183.9779.7983.3380.5283.4380.3784.1882.12
NVIDIA Nemotron 3 Nano Omni 30B A3B (Open)67.6583.2165.8073.5383.7780.5774.8174.1874.52
Qwen3-Omni-30B-A3B-Thinking (Open)65.9780.9265.8071.5785.0680.5776.0470.7373.60
Gemini 2.5 Flash (Closed)73.8266.4172.0468.0378.6781.8769.8677.0973.06
Qwen3-Omni-30B-A3B-Instruct (Open)66.8180.9264.7766.3481.1781.1471.8771.8271.85
Gemini 2.0 Flash (Closed)62.1873.2863.7363.7276.6275.4367.2368.5567.84
Qwen2.5-Omni-7B-Instruct (Open)48.3269.4758.5558.1776.6273.1464.6159.0962.07
Gemini 2.5 Flash Lite (Closed)57.5668.7056.4852.6179.2269.7163.5260.0061.90
Daily-Omni-Baseline-Qwen2.5 (Open)51.6868.7060.1053.9278.5771.4363.9959.2761.82
Gemini 2.0 Flash Lite (Closed)55.0464.8958.0354.2574.0372.0062.4460.0061.32
Qwen2.5-Omni-3B-Instruct (Open)50.8469.4753.8953.9275.9770.2962.6057.4560.23
Ola (7B) (Open)40.3461.0740.4143.4663.6469.7151.4749.8250.71
VideoLLaMA2 (7B) (Open)35.7135.8835.7531.7040.9134.2938.0231.8235.17
Unified-IO-2 XL (3B) (Open)30.2530.5325.3929.0833.1221.7128.1328.5528.32
Unified-IO-2 XXL (8B) (Open)25.6331.3026.4225.8235.0629.7126.7430.0028.24
Unified-IO-2 L (1B) (Open)27.3122.9026.4227.7829.8729.1427.6727.0927.40

† Qwen3.5-Omni-Plus returned valid answers for 1,175 of 1,197 samples. The remaining 22 samples could not be evaluated after repeated API attempts because of request-size limits (13) and content filtering (9). Its scores use the valid API responses in each subset as the denominator.

Omni-Modal Language Models (Visual Only)

MethodsAV AlignComp.Ctx. Und.Evt. Seq.Infer.Reas.30s60sAvg
Qwen3-Omni-30B-A3B-Instruct (Open)47.9067.1856.4853.9270.7861.1457.9657.6457.81-14.0
Qwen3-Omni-30B-A3B-Thinking (Open)44.9664.8955.9659.4867.5360.0056.2659.4557.73-15.9
Gemini 2.0 Flash (Closed)39.0864.1256.4856.2167.5362.2956.5755.4556.06-11.8
Gemini 2.0 Flash Lite (Closed)43.7058.0253.8945.1064.2960.5753.0151.6452.38-8.9
Qwen2.5-Omni-7B-Instruct (Open)34.4558.7847.6749.6762.9954.8648.6951.0949.79-12.3
Qwen2.5-Omni-3B-Instruct (Open)37.3951.9144.5641.1864.2948.0046.5245.6446.12-14.1
Gemini 2.5 Flash (Closed)37.5537.2140.4344.7857.0553.2942.3547.4644.61-28.5
Gemini 2.5 Flash Lite (Closed)36.9745.8037.3139.5459.7447.4344.6741.2743.11-18.8

Omni-Modal Language Models (Audio Only)

MethodsAV AlignComp.Ctx. Und.Evt. Seq.Infer.Reas.30s60sAvg
Qwen3-Omni-30B-A3B-Instruct (Open)54.2069.4751.8151.6374.0378.8663.3758.1860.99-10.9
Qwen3-Omni-30B-A3B-Thinking (Open)54.6267.9449.2251.3177.2777.7165.2255.2760.65-13.0
Gemini 2.5 Flash (Closed)46.6455.7344.5642.4870.7878.8655.6452.1854.05-19.0
Gemini 2.5 Flash Lite (Closed)42.0261.8341.9745.1068.8365.1454.2548.9151.80-10.1

Visual Language Models (Visual Only)

MethodsAV AlignComp.Ctx. Und.Evt. Seq.Infer.Reas.30s60sAvg
Qwen3-VL-30B-A3B-Instruct (Open)47.4868.7052.3355.8867.5361.1457.3457.2757.31
Qwen3-VL-8B-Instruct (Open)44.5463.3650.7859.8069.4858.8656.4157.2756.81
GPT-4o (Closed)47.9062.6052.3352.6166.2366.2955.6457.4556.47
Qwen3-VL-4B-Instruct (Open)43.7061.0754.4053.2768.1858.8654.4056.0055.14
Qwen2.5-VL-7B-Instruct (Open)36.9746.5633.6837.9151.9544.0039.2642.3640.68
Qwen2.5-VL-3B-Instruct (Open)35.7143.5134.7233.6643.5139.4337.7137.0937.43

Audio Language Models (Audio Only)

MethodsAV AlignComp.Ctx. Und.Evt. Seq.Infer.Reas.30s60sAvg
Audio Flamingo 3 (7B) (Open)40.7655.7343.0140.5265.5868.0050.2349.4549.87
Qwen2-Audio (7B) (Open)28.9935.8827.4632.0333.7733.1431.2231.8231.50

Textual Language Models (Without Visual and Audio)

MethodsAV AlignComp.Ctx. Und.Evt. Seq.Infer.Reas.30s60sAvg
GPT-4o (Closed)33.1943.5128.5030.3944.8146.8636.4836.1836.34
Deepseek-V3 (671B) (Open)31.9341.2229.0229.4144.8146.2935.2436.0035.59
Qwen2.5-14B-Instruct (Open)30.2539.6927.9828.4342.2142.8632.1535.8233.83

Requirements

To install requirements:

pip install -r requirements.txt

To test our benchmark, you should download Videos.tar and 'qa.json' from Huggingface to this directory and extract the Videos/ folder and qa.json to this directory.

QA Generation

QA Generation Pipeline

πŸ“‹ We provide script to reproduce our Daily-Omni QA Generation pipeline. Run the following command to generate QA pairs. To run the script, first you should revise the config.py file to set the parameters:

  • Set the api keys, base_urls and model_name. You can create api keys from the following links: Gemini, OpenAI, Deepseek, Aliyun

  • Set BASE_DIR and CSV_PATH to the Video Folder you want to annotate and the path to the csv file that records the videos. You can use example_videos and example_metadata.csv as templetes.

  • Set MAX_WORKERS_PROCESSES to the number of processes you want to use to run the pipeline. You can set execution_mode in run_pipeline.py to choose the execution mode if your api key has parallel requests limitation.

  • Set run_pipeline_flags in run_pipeline.py to choose which part of the pipeline you want to run.

Test the pipeline with:

python run_pipeline.py

Test Daily-Omni

Model with api

πŸ“‹To test the benchmark on third party models(Gemini, GPT-4o, Deepseek) with api and reproduce the results, you can use the script provided in test_model_api/

python test_model_api/test_model.py --model <model_name> --mode <Execution_mode> --max_items <Maximum number of QA items to process (for testing)>

You can check the model options in test_model_api/test_config.py

Model running locally

πŸ“‹To test the benchmark on third party models(Qwen2.5-Omni, Qwen2.5-VL, VideoLLaMA2, Ola, Unified-IO 2) with local machines, check the code in test_model/

All local test_model/*/testmodel.py scripts now use a unified modality argument:

--input_mode {all,visual,audio}

The default is --input_mode all.

  • all: video + audio
  • visual: video only
  • audio: audio only

--use_audio_in_video has been removed. For scripts that save per-item JSONL results, raw model output is saved by default.

Qwen2.5-Omni

You should install the dependencies with instructions from the official Qwen2.5-Omni repo

Run the test script with the following command:

python test_model/Qwen2.5-Omni/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH --model_name_or_path MODEL_NAME_OR_PATH --processor_name_or_path PROCESSOR_NAME_OR_PATH --input_mode all

You can switch modality with --input_mode all|visual|audio.

Qwen2.5-VL

You should install the dependencies with instructions from the official Qwen2.5-VL repo

Run the test script with the following command:

python test_model/Qwen2.5-VL/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH --model_name_or_path MODEL_NAME_OR_PATH --processor_name_or_path PROCESSOR_NAME_OR_PATH --input_mode all

Qwen2.5-VL is visual-only. --input_mode all will fall back to visual, and audio is not supported.

Qwen3-VL

You should install the dependencies with instructions from the official Qwen3-VL repo

Run the test script with the following command:

python test_model/Qwen3-VL/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH --model_name_or_path MODEL_NAME_OR_PATH --input_mode all --batch_size 8

Qwen3-VL is visual-only. --input_mode all will fall back to visual, and audio is not supported.

VideoLLaMA2

You should install the dependencies with instructions from the official VideoLLaMA2 repo

Run the test script with the following command:

python test_model/VideoLLaMA2-av/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH --input_mode all

Unified-IO 2

You should install the dependencies with instructions from the official Unified-IO 2 repo

Run the test script with the following command:

python test_model/unified-io-2.pytorch/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH --model MODEL_NAME --input_mode all

unified-io-2.pytorch currently only supports --input_mode all.

Ola

You should install the dependencies with instructions from the official Ola repo

Run the test script with the following command:

python test_model/Ola/inference/testmodel.py --video_base_dir VIDEO_BASE_DIR --json_file_path JSON_FILE_PATH

Ola currently uses its own script interface and is not aligned to the unified --input_mode CLI above.

Subsampling stability test

Use run_subsampling_stability.sh to run subsampling stability evaluation for the current model set.

./run_subsampling_stability.sh

What it does:

  • Runs per-item evaluation first and saves *_items.jsonl for each model.
  • Samples 20,40,60,80,100% of questions by default.
  • Uses all QA belonging to the sampled videos to compute accuracy.
  • Repeats the experiment R times and reports mean Β± std and median [p5, p95].

Default model set:

  • Qwen/Qwen2.5-Omni-7B-Instruct
  • Qwen/Qwen3-Omni-30B-A3B-Instruct
  • gemini-2.5-flash
  • gemini-2.5-flash-lite

Useful environment variables:

  • QWEN_INPUT_MODE / QWEN25_INPUT_MODE: modality, default all
  • STABILITY_REPEATS: repeat count, default 200
  • SUBSAMPLE_PCTS: comma-separated percentages, default 20,40,60,80,100
  • INCLUDE_QWEN_THINKING=1: include Qwen3-Omni-30B-A3B-Thinking
  • GEMINI_API_KEY_1 and GEMINI_API_KEY_2: Gemini API keys

Outputs are written to runs/subsampling_stability_<timestamp>/, including per-item JSONL, stability_repeat_results.csv, stability_summary.csv, and stability_summary.txt.

Test Daily-Omni Agent

Daily-Omni Agent

We provide a script to test Daily-Omni Agent on Daily-Omni benchmark. For efficiency, we used the API provided by Bailian Aliyun for Qwen2.5-VL and Qwen2.5, while Qwen2-Audio was deployed locally. According to the official documentation, qwen2.5-vl-7b-instruct and qwen2.5-14b-instruct provided by Bailian Aliyun are identical to their open-source counterparts. Our code implements direct passing of local_video_path to the Qwen2.5-VL API. However, this functionality might require you to contact Aliyun customer service to enable. If direct path input is not activated, you can alternatively pass a list of video frames, though this may result in suboptimal performance.

To run the agent, you need to setup a new environment for Qwen2-Audio according to the instructions in the official repository.

Then, launch the Qwen2-Audio server locally with running:

python baseline/qwen_audio.py

Segment the video and audio clips:

python baseline/segment_av.py

Run Daily-Omni Agent on Daily-Omni benchmark with the following command:

python baseline/base_model.py

This script will automatically evaluate the performance of the model on the Daily-Omni benchmark.

Citation

@misc{zhou2026dailyomniaudiovisualreasoningtemporal,
      title={Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities}, 
      author={Ziwei Zhou and Rui Wang and Zuxuan Wu and Yu-Gang Jiang},
      year={2026},
      eprint={2505.17862},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2505.17862}, 
}