Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation
July 22, 2025 ยท View on GitHub
๐ฅ News
2025.05.16๐ COVER has been accepted by ACL 2025 findings.
๐ COVER Overview
We introduce COVER , a multidimensional multimodal benchmark that systematically evaluates MLLMs across the abstract-concrete and perception-cognition dimensions. Our benchmark includes approximately 2,800 videos, which are paired with around 12,000 to 13,000 individual QA instances. As stated in figure, the enhanced version of our dataset consists of about 2.9k question pairs, with each pair comprising at least three individual QA items:
- One original question
- One counterfactual question
- At least one sub-question (often multiple)
๐ Dataset Examples
Click to expand more examples
๐ Dataset
Download COVER data here
License:
COVER is only used for academic research. Commercial use in any form is prohibited.
The copyright of all videos belongs to the video owners.
If there is any infringement in COVER, please email zhouqiji@westlake.edu.cn and we will remove it immediately.
Without prior approval, you cannot distribute, publish, copy, disseminate, or modify COVER in whole or in part.
You must strictly comply with the above restrictions.
๐ Experimental Results
- General assessment results of COVER.
- Evaluation results of different MLLMs on our quadrant formulation.
- Comparison between CoT and Guide-CoT performance across MLLMs on the COVER.
- Heatmaps of task performance for Gemini-1.5-pro and InternVL2.5-78B.
๐ Citation
If you find our work helpful for your research, please consider citing our work.
@article{zhou2025reasoning,
title={Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation},
author={Zhou, Qiji and Gong, Yifan and Bao, Guangsheng and Qiu, Hongjie and Li, Jinqiang and Zhu, Xiangrong and Zhang, Huajian and Zhang, Yue},
journal={arXiv preprint arXiv:2503.10691},
year={2025}
}