README.md

August 12, 2025 · View on GitHub

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

If you find this project useful, please give us a star🌟.

News

  • May 27, 2025. We have released our paper in the arxiv. Data and model will be released soon.

Introduction

Unlike existing LMMs that rely on shortcut learning or chained inference, Griffon-R introduces a unified, single-pass mechanism that mimics human-like understanding-thinking-answering. It bridges foundational visual capabilities with high-level reasoning to produce faithful, traceable answers — no external tools or multiple passes required. We also release 334K curated visual instruction samples spanning diverse scenes and reasoning challenges, fueling Griffon-R’s strong generalization across benchmarks like VSR, CLEVR, MMBench, and ScienceQA.

Demo