README.md
August 12, 2025 · View on GitHub
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
If you find this project useful, please give us a star🌟.
News
-
May 27, 2025.We have released our paper in the arxiv. Data and model will be released soon.
Introduction
Unlike existing LMMs that rely on shortcut learning or chained inference, Griffon-R introduces a unified, single-pass mechanism that mimics human-like understanding-thinking-answering. It bridges foundational visual capabilities with high-level reasoning to produce faithful, traceable answers — no external tools or multiple passes required. We also release 334K curated visual instruction samples spanning diverse scenes and reasoning challenges, fueling Griffon-R’s strong generalization across benchmarks like VSR, CLEVR, MMBench, and ScienceQA.

Demo
