SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?

February 5, 2026 · View on GitHub

ICLR 2026 Project-Website arxiv Kaggle GitHub HuggingFace HuggingFace

Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez

Computational Intelligence and Operations Laboratory (CIOL) • Shahjalal University of Science and Technology (SUST) • Monash University • Qatar Computing Research Institute (QCRI)

The Fourteenth International Conference on Learning Representations (ICLR 2026)


SpatiaLab is a benchmark for evaluating spatial reasoning in vision–language models (VLMs) under real-world, in-the-wild visual conditions.
It includes 1,400 visual question–answer pairs across 6 core spatial categories and 30 subcategories, supporting both multiple-choice (MCQ) and open-ended evaluation formats.
SpatiaLab exposes substantial gaps between state-of-the-art VLMs and human performance.

SpatiaLab overview figure


Overview

Spatial reasoning is fundamental to human intelligence and real-world embodied AI.
SpatiaLab provides a comprehensive evaluation suite (1,400 QA pairs) across six core spatial categories:

  • Relative Positioning
  • Depth & Occlusion
  • Orientation
  • Size & Scale
  • Spatial Navigation
  • 3D Geometry

It is designed to test VLMs in realistic, unconstrained scenes and highlights large performance gaps between models and humans.


Benchmark Structure and Categorization

SpatiaLab comprises 1,400 validated QA items organized into 6 main categories and 30 subcategories (5 subcategories each).

CategoryExample sub-tasks (5 each)
Relative PositioningLeft/Right, Above/Below, Between, Adjacency, Corner/Angle
Depth & OcclusionPartial occlusion, Complete occlusion, Layer order, Reflection/visibility, Hidden feature
OrientationRotation angle, Facing, Tilt, Tool handedness, Mirror
Size & ScaleRelative size, Scale ratio, Big/Small, Proportion, Size consistency
Spatial NavigationPath existence, Obstacle avoidance, Turn sequence, Viewpoint visibility, Accessibility
3D Geometry3D containment, Intersection, Volume ordering, Pose matching, Stability

Category distribution pie chart


Data Collection and Annotation

SpatiaLab images and QA pairs were created through:

  • Web crawling
  • Targeted retrieval
  • Manual snapshots

A 3-stage annotation and verification process yields the final 1,400 QA items.

Data pipeline


Task Formats & Sample Instances

Every subcategory includes both:

  • Multiple-choice (MCQ) items (discriminative accuracy)
  • Open-ended items (generative spatial reasoning)

Example tasks

Annotation & QC

Annotators were trained; items passed a 3-stage review to produce gold-standard annotations.
Inter-annotator reliability and judge agreement are reported in the paper (appendix).


Results

Tables below are transcribed from the paper.

Multiple Choice Evaluation Accuracy (%) on SpatiaLab-MCQ

Model3D Geom. (#238)Dep. & Occu. (#259)Orientation (#202)Relat. Posit. (#212)Size & Scale (#252)Spati. Navig. (#237)Overall (#1400)
Random Choice25.0025.0025.0025.0025.0025.0025.00
GPT-4o-mini47.0639.0047.0347.1749.6049.7946.50
GPT-5-mini48.7454.8360.4062.7444.8456.5454.29
Gemini-2.0-Flash47.0655.2153.9658.0254.3746.8452.50
Gemini-2.5-Flash44.9648.2648.0256.1342.4651.0548.29
Claude 3.5 Haiku42.4442.0846.5346.2335.7145.9942.93
Mistral Medium 3.146.6449.8147.5261.7941.6741.7747.93
InternVL3.5-1B33.6132.4323.2737.2631.7530.8031.64
InternVL3.5-2B34.0331.6631.6840.5732.5432.4933.71
Qwen-VL2.5-3B-Instruct41.1835.5246.0440.0947.2239.2441.43
InternVL3.5-4B42.8642.8642.0854.7236.5142.1943.29
Gemma-3-4B-it43.7034.3646.5345.7537.3037.9740.57
Qwen-VL2.5-7B-Instruct42.8637.8442.5746.2342.0635.4441.00
Llama-3.2-11B-Vision-Instruct26.4730.5020.3042.9230.5632.0730.50
Gemma-3-27B-it43.2840.1548.0254.2548.0247.2646.57
Qwen-VL2.5-32B-Instruct41.1840.1546.5345.2845.2441.7743.21
InternVL3.5-72B50.0057.1453.4766.0449.2154.8554.93
Qwen-VL2.5-72B-Instruct47.0648.6551.9854.2543.6548.9548.86
Llama-3.2-90B-Vision-Instruct46.2252.1250.5058.9646.8348.5250.36
o4-mini-medium51.2658.3054.9564.1540.8751.4853.21
Gemini-2-Flash-Thinking37.8241.3141.5845.7550.4043.0443.36
Gemini-2.5-Flash-Thinking45.8053.6752.9756.6055.1653.5952.93
SpaceOm42.4438.6148.0237.7442.8639.2441.36
SpaceThinker-Qwen2.5VL-3B40.3437.8447.0338.2143.2537.9740.64
SpaceQwen2.5-VL-3B-Instruct31.5135.1437.6237.7450.7947.2640.14
Human Baseline93.7074.1391.5891.5188.8987.7687.57

Open-ended Evaluation Accuracy (%) on SpatiaLab-OPEN

Model3D Geom. (#238)Dep. & Occu. (#259)Orientation (#202)Relat. Posit. (#212)Size & Scale (#252)Spati. Navig. (#237)Overall (#1400)
GPT-4o-mini23.5316.6023.2730.6617.8621.9426.00
GPT-5-mini45.3834.7537.1349.5342.4637.1340.93
Gemini-2.0-Flash31.9324.3227.2331.1326.1924.4727.43
Gemini-2.5-Flash34.0326.6431.6838.6826.5929.5430.93
Claude 3.5 Haiku26.0518.9224.7525.9420.2421.1022.64
Mistral Medium 3.125.2119.3121.7829.2515.0816.8821.00
InternVL3.5-1B05.8809.6509.9013.6809.1310.1309.64
InternVL3.5-2B12.1811.2010.8923.5811.9018.1414.50
Qwen-VL2.5-3B-Instruct15.5508.4915.3510.8518.2509.2812.93
InternVL3.5-4B19.3317.7615.8419.8116.2718.9918.00
Gemma-3-4B-it20.1713.1314.8523.5815.0819.8317.64
Qwen-VL2.5-7B-Instruct15.1315.8320.3027.8315.8719.8318.86
Llama-3.2-11B-Vision-Instruct16.8116.9922.2825.0013.4918.5718.57
Gemma-3-27B-it22.6916.2224.7534.4322.6221.9423.43
InternVL3.5-72B22.6920.4620.3031.6019.8426.1623.36
Qwen-VL2.5-72B-Instruct26.8920.8525.2530.6624.6020.6824.64
Llama-3.2-90B-Vision-Instruct22.6923.1721.2928.3021.8327.0024.00
GLM-4.5V-106B-MoE31.0920.4625.2526.4224.2124.4725.21
o4-mini-medium40.7632.8232.1842.9244.0534.1837.86
Gemini-2-Flash-Thinking31.0927.4131.1934.4329.3729.5430.36
Gemini-2.5-Flash-Thinking37.1445.4536.3637.1421.7422.2232.77
SpaceOm12.6106.9515.8411.7918.6512.2412.93
SpaceThinker-Qwen2.5VL-3B13.4509.2717.8210.3819.4410.1313.36
SpaceQwen2.5-VL-3B-Instruct12.6103.8613.8609.4311.9011.3910.36
Human Baseline73.5350.1970.3069.8165.4862.8764.93

Key Takeaways

  • Large human–model gap.
    MCQ: top models ~55% vs humans 87.6%.
    Open-ended: best ~41% vs humans ~65%.

  • Open-ended is much harder.
    Average MCQ → Open-ended drop is substantial across models.

  • Scale alone is not sufficient.
    Some large models remain weak; small models often cluster near the bottom.

  • Spatial “specialists” don’t necessarily generalize.
    Specialized spatial models can underperform broadly, especially in open-ended settings.


Error Analysis Summary

Common failure modes observed across models:

  • Spatial mislocalization in cluttered scenes (wrong referents)
  • Perspective/scale mistakes (over-reliance on size priors)
  • Occlusion and ordering failures (thin/partially hidden structures)
  • Fluent but visually ungrounded open-ended answers
  • Multi-cue integration failures (depth + size + ordering)
  • Poor confidence calibration in open-ended generation

Methods

  • Image sources: web crawling, targeted retrieval, manual capture
  • Annotation: trained annotators + 3-stage review/QC
  • Evaluation:
    • MCQ: option selection + exact match
    • Open-ended: free-form generation + judge scoring (validated against human agreement)
  • Metrics: accuracy + agreement measures (e.g., Cohen’s / Fleiss’ kappa reported in paper)

Performance Improvement Approaches (Explored)

  • Inherent reasoning-enabled models
  • Chain-of-Thought (CoT) prompting
  • CoT + self-reflection
  • Supervised fine-tuning (SFT)
  • Multi-agent system (SpatioXolver)

Citation

@inproceedings{
wasi2026spatialab,
title={SpatiaLab: Can Vision{\textendash}Language Models Perform Spatial Reasoning in the Wild?},
author={Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=fWWUPOb0CT}
}