Awesome-MLLM-Reasoning-Collection

July 31, 2026 · View on GitHub

License: MIT arXiv

⚠️ Repository notice

On 31 May 2026, this repository was targeted by an apparent coordinated fake-star attack that also affected many other open-source repositories. Its star count rose abnormally from approximately 673 to more than 14,000 within a single day, despite no promotion or involvement from the maintainers. We reported the incident to GitHub but have not received a substantive response. We remain sincerely grateful to the 670+ genuine supporters whose earlier stars may no longer be displayed—your support and trust have always been deeply appreciated. 💙

👏 Welcome to the Awesome-MLLM-Reasoning-Collections repository! This repository is a carefully curated collection of papers, code, datasets, benchmarks, and resources focused on reasoning within Multimodal Large Language Models (MLLMs).

Feel free to ⭐ star and fork this repository to keep up with the latest advancements and contribute to the community.

alt text

A conceptual trajectory of multimodal reasoning, evolving from static image-level understanding, through temporal video and audio reasoning, to holistic omni-level reasoning, and finally toward embodied embedding reasoning with perception–action interaction. This progression reflects increasing reasoning scope, compositionality, and interactivity.

Citation

If you find this repository or our survey useful for your research, please consider citing:

@article{hu2026static,
  title   = {From Static Perception to Interactive Decision: A Survey of Multimodal Reasoning},
  author  = {Hu, Jian and Cheng, Zixu and Ma, Yinghao and Dixit, Satvik and Pan, Bikang and Chen, Lei and Ma, Lin and Zeng, Zhixiong and Wang, Jiangya and Benetos, Emmanouil and others},
  journal = {researchgate preprint},
  year    = {2026}
}

Table of Contents

Papers and Projects 📄

alt text

An evolutionary landscape of several representative multi-modal reasoning frameworks from 2022 to 2025.

Commonsense Reasoning

Image MLLM

Video MLLM

Audio MLLM

  • Utilizing GRPO to enhance audio reasoning performance

Omni MLLM

Reasoning Segmentation and Detection

Image MLLM

Video MLLM

Audio MLLM

Omni MLLM

Spatial and Temporal Grounding and Understanding

Image MLLM

Video MLLM

Audio MLLM

Omni MLLM

Math Reasoning

Image MLLM

Chart Rasoning

Benchmark

Visual-Audio Generation

Image MLLM

Video MLLM

Audio MLLM

Reasoning with Agent/Tool

Medical Reasoning

Image MLLM

Audio MLLM

Omni MLLM

Embodied Reasoning

Others

Image MLLM

Video MLLM

Audio MLLM

Omni LLM

Benchmarks 📊

DateProjectTaskLinks
26.03MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image ReasoningMulti-image Reasoning[📑 Paper] [🌐 Project]
26.03RIVER: A Benchmark for Real-World Video Reasoning in Long and Short ContextsVideo Temporal Reasoning[📑 Paper]
26.03UniG2U-Bench: Comprehensive Benchmark for Unified Generation and Understanding MLLMsMultimodal Evaluation[📑 Paper]
26.03AgentVista: Generalizable Multi-Task Agent with Diverse Visual ManipulationMulti-Task Agent[📑 Paper]
26.02A Very Big Video Reasoning Suite (VBVR): 1M+ video clips across 200 reasoning tasksVideo Reasoning[📑 Paper] [🤗 Model] [🤗 Data]
26.02OmniGAIA: Omni-Modal AI Agent Benchmark with hindsight-guided explorationOmni-Modal Agent Reasoning[📑 Paper] [💻 Code] [🤗 Data]
26.02SpatiaLab: Wild Spatial Reasoning benchmark across 6 VQA categoriesSpatial Reasoning[📑 Paper] [💻 Code] [🤗 Data]
26.02MuRGAt: Multimodal Fact-Level Attribution benchmark for verifiable reasoningMultimodal Attribution[📑 Paper] [💻 Code]
26.02DeepVision-103K: Verifiable multimodal math dataset for RLVR trainingMath Reasoning[📑 Paper] [💻 Code] [🤗 Data]
26.02UniVBench: Unified evaluation for video foundation models across understanding, generation, editingVideo Foundation Model Evaluation[📑 Paper] [💻 Code]
26.02RISE-Video: Benchmark for video generators decoding implicit world rulesVideo Generation Reasoning[📑 Paper] [💻 Code] [🤗 Data]
26.02SAW-Bench: Egocentric Situated Awareness evaluation with 786 smart-glass videos and 2,071+ QA pairsSpatial Reasoning[📑 Paper]
26.02BrowseComp-V3: 300-question visual benchmark for complex multi-hop multimodal web searchMultimodal Browsing[📑 Paper]
26.02BiManiBench: Hierarchical benchmark for bimanual coordination evaluation in MLLMsBimanual Robotics[📑 Paper] [💻 Code]
26.01MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric MethodsMultimodal Reasoning[📑 Paper] [🤗 Model] [🤗 Data]
26.01ChartVerse: Scaling Chart Reasoning via Reliable Programmatic SynthesisChart Reasoning[📑 Paper] [💻 Code] [🤗 Model] [🤗 Data]
26.01VideoLoom: Joint Spatial-Temporal Understanding with LoomBenchSpatial-Temporal Reasoning[📑 Paper] [💻 Code] [🤗 Model]
26.01PROGRESSLM: Towards Progress Reasoning in Vision-Language ModelsTask Progress Reasoning[📑 Paper] [💻 Code] [🤗 Data]
26.01FutureOmni: Evaluating Future Forecasting from Omni-Modal ContextOmni-Modal Temporal Reasoning[📑 Paper]
26.01Afri-MCQA: Multimodal Cultural Question Answering for African LanguagesMultilingual Multimodal Reasoning[📑 Paper]
26.01AVMeme Exam: A Multimodal Multilingual Multicultural BenchmarkCultural Multimodal Reasoning[📑 Paper]
25.12HERBench: Multi-Evidence Integration in Video Question AnsweringVideo Reasoning[📑 Paper]
25.12SVBench: Evaluation of Video Generation Models on Social ReasoningVideo Social Reasoning[📑 Paper]
25.12IF-Bench: Benchmarking MLLMs for Infrared ImagesInfrared Image Understanding[📑 Paper]
25.12VABench: Comprehensive Benchmark for Audio-Video GenerationAudio-Video Generation[📑 Paper]
25.11MME-CC: Challenging Multi-Modal Evaluation Benchmark of Cognitive CapacityCognitive Capacity[📑 Paper]
25.11GGBench: Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsGeometric Reasoning[📑 Paper]
25.11WEAVE: Benchmarking In-context Interleaved Comprehension and GenerationMultimodal Comprehension & Generation[📑 Paper]
25.10Uni-MMMU: Massive Multi-discipline Multimodal Unified BenchmarkMultimodal Multi-discipline Reasoning[📑 Paper]
25.10PhysToolBench: Benchmarking Physical Tool Understanding for MLLMsPhysical Tool Understanding[📑 Paper]
25.10BEAR: Benchmarking Multimodal Language Models for Atomic Embodied CapabilitiesEmbodied AI Capabilities[📑 Paper]
25.10OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsLong-context, Video-Audio Unerstanding & Reasonin[📑 Paper] [💻 Code] [🌐 Project] [🤗 Data]
25.10XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsCapability Balancing among Different Modalities[📑 Paper] [💻 Code] [🌐 Project]
25.10StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQATermporal Reasoning[📑 Paper]
25.10Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and BenchmarkCommon Sense Omni Reasoning[📑 Paper]
25.09MARS2 2025 Challenge on Multimodal ReasoningMultimodal Reasoning Challenge[📑 Paper]
25.09Visual-TableQA: Open-Domain Benchmark for Reasoning over Table ImagesTable Reasoning[📑 Paper]
25.09AHELM: A Holistic Evaluation of Audio-Language ModelsAudio-Language Understanding[📑 Paper]
25.09MDAR: A Multi-scene Dynamic Audio Reasoning BenchmarkComplex, Multi-scene, & Dynamically Evolving Speech & Audio Reasonin[📑 Paper] [💻 Code]
25.09MiMo-Audio-Eval ToolkitSpeech/Sound/Music Reasoning[💻 Code]
25.08SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language ModelsSpeech Reasoning[📑 Paper] [💻 Code] [Data]
25.08MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceLong-form, Spatial, and Multi-audio Reasoning on Speech/Music/Sound[📑 Paper] [🤗 Data]
25.08R²-AVSBench: Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationSegmentation Reasoning[📑 Paper] [🤗 Data]
25.07Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and UnderstandingVideo Reasoning and Understanding[📑 Paper]. [🌐 Project] [🤗 Data]
25.06FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning EvaluationFinancial Multi-Modal Reasoning Reasoning[📑 Paper]. [💻 Code]. [🤗 Data]
25.06MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosVideo Reasoning[📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data]
25.06OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsSpatial Reasoning[📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data]
25.06MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkPhonatics, Prosody, Rhetoric, Syntactics, Semantics, and Paralinguistics in Speech Understanding & Reasoning[📑 Paper] [💻 Code] [🤗 Data]
25.05Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across ModalitiesVideo&Audio Reasoning[📑 Paper] [💻 Code] [🌐 Project] [🤗 Data]
25.05MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their MixMulti-step Audio Reasoning[📑 Paper]. [💻 Code]. [🎥 demo] [🤗 Data]
25.05On Path to Multimodal Generalist: General-Level and General-BenchMultimodal Generation[🌐 Project] [📑 Paper] [🤗 Data]
25.04VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsVisual Reasoning[🌐 Project] [📑 Paper] [💻 Code] [🤗 Data]
25.04IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsImage-Grounded Video Perception and Reasoning[📑 Paper] [💻 Code]
25.04Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual EditingReasoning-Informed viSual Editing[📑 Paper] [💻 Code]
25.04CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction FollowingMusic Information Retrieval & Knowledge[📑 Paper] [💻 Code]
25.03MAVERIX: Multimodal Audio-Visual Evaluation Reasoning IndeXCommon Sense Omni Reasoning[📑 Paper] [🌐 Project]
25.03V-STaR : Benchmarking Video-LLMs on Video Spatio-Temporal ReasoningSpatio-temporal Reasoning[🌐 Project] [📑 Paper] [🤗 Data]
25.03MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsSpatio-temporal Understanding[📑Paper]
25.03Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning3D-CoT[📑 Paper] [🤗 Data]
25.02MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal ModelsMM-IQ[📑 Paper] [💻 Code]
25.02MM-RLHF: The Next Step Forward in Multimodal LLM AlignmentMM-RLHF-RewardBench, MM-RLHF-SafetyBench[📑 Paper]
25.02ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal ModelsZeroBench[🌐 Project] [🤗 Dataset] [💻 Code]
25.02MME-CoT: Benchmarking Chain-of-Thought in LMMs for Reasoning Quality, Robustness, and EfficiencyMME-CoT[📑 Paper] [💻 Code]
25.02OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceMM-AlignBench[📑 Paper] [💻 Code]
25.01AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMsAdversarial attack, Compositional reasoning, and Modality-specific dependency in Visual&Audio[📑 Paper]
25.01LlamaV-o1: Rethinking Step-By-Step Visual Reasoning in LLMsVRCBench[📑 Paper] [💻 Code]
24.12Online Video Understanding: A Comprehensive Benchmark and Memory-Augmented MethodVideoChat-Online[Paper📑] [Code💻]
24.11VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsVLRewardBench[📑 Paper]
24.11Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosMH-VidQA[Paper📑] [Code💻]
24.10OmnixR: Evaluating Omni-modality Language Models on Reasoning across ModalitiesVideo&Audio Reasoning[📑 Paper]
24.10MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkAudio Understanding & Reasoning[🌐 Project] [📑 Paper] [💻Code] [🤗 Data]
24.09MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningVideo Causal Reasoning[📑 Paper] [💻Code] [🤗 Data]
24.09OmniBench: Towards The Future of Universal Omni-Language ModelsReasoning with Image & Speech/Sound/Music[📑 Paper] [Code💻] [🌐 Project] [🤗 Data]
24.08MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language ModelsMusic Knowledge & Reasoning[🌐 Project] [📑 Paper] [💻Code] [ Data]
24.07REXTIME: A Benchmark Suite for Reasoning-Across-Time in VideosREXTIME[Paper📑] [Code💻]
24.06AudioBench: A Universal Benchmark for Audio Large Language ModelsSpeech & Sound Understanding[Paper📑] [Code🖥️]
24.06ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code GenerationChartBench[Project🌐] [Paper📑] [Code🖥️]
24.05M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-ThoughtM3CoT[📑 Paper]
24.02AIR-Bench: Benchmarking Large Audio-Language Models via Generative ComprehensionSpeech & Sound Understanding[📑 Paper] [Code💻]
23.10CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsAudio Reasoning (Attributes & Orders)[Project🌐] [Paper📑]

Open-source Projects

ProjectGitHub StarsLinks
Reason-RFTReason-RFT💻 GitHub 🤗 Dataset
EasyR1EasyR1💻 GitHub
Multimodal Open R1Multimodal Open R1💻 GitHub 🤗 Model 🤗 Dataset
LMM-R1LMM-R1💻 GitHub
MMR1MMR1💻 GitHub 🤗 Model 🤗 Dataset
R1-VR1-V💻 GitHub 🎯 Blog 🤗 Dataset
R1-Multimodal-JourneyR1-Multimodal-Journey💻 GitHub
VLM-R1VLM-R1💻 GitHub 🤗 Model 🤗 Dataset 🤗 Demo
R1-VisionR1-Vision💻 GitHub 🤗 Cold-Start Dataset
R1-OnevisionR1-Onevision💻 GitHub 🤗 Model 🤗 Dataset 🤗 Demo 📝 Report
Open R1 VideoOpen R1 Video💻 GitHub 🤗 Model 🤗 Dataset
Video-R1Video-R1💻 GitHub 🤗 Dataset
Open-LLaVA-Video-R1Open-LLaVA-Video-R1💻 GitHub
R1V-FreeR1V-Free💻 GitHub
SeekWorldSeekWorld💻 GitHub
IE-Critic-R1SeekWorld💻 GitHub
🤗 Model
🤗 Data
🤗 ColdStart SFT

Contributing

If you are interested in contributing, please refer to HERE for instructions in contribution.