GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

July 1, 2025 Β· View on GitHub

πŸŽ‰ Exciting News:
Our paper has been ACCEPTED to ICCV 2025 – International Conference on Computer Vision!
in Honolulu, Hawaii!

Oryx Video-ChatGPT

Muhammad Sohail Danish*, Muhammad Akhtar Munir*, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro , Alexandre Lacoste and Salman Khan

* Equally contributing first authors

Mohamed bin Zayed University of AI, University College London, LinkΓΆping University, IBM Research Europe, UK, ServiceNow Research, Australian National University

paper Website HuggingFace

Official GitHub repository for GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks.

πŸ“’ Latest Updates

  • Jun-30-25: Code is released.
  • Apr-04-25: We have evaluated recent geospatial Vision-Language Models (VLMs).
  • Dec-02-24: We release the benchmark dataset huggingface link.
  • Dec-02-24: Arxiv Preprint is released arxiv link. πŸ”₯πŸ”₯

πŸ› οΈ Leaderboard Coming Soon!

The leaderboard will be released shortly. Follow this repository for updates!


πŸ’‘ Overview

Figure: GEOBench-VLM comprehensively covers 31 fine-grained tasks categorized into 8 broad categories: scene and object classification, object detection, segmentation, captioning, event detection, non-optical and temporal understanding tasks.

Abstract: While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for applications such as environmental monitoring, urban planning, and disaster management. Key challenges in the geospatial domain include temporal change detection, large-scale object counting, tiny object detection, and understanding relationships between entities in remote sensing imagery. To bridge this gap, we present GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, segmentation, and temporal analysis. Our benchmark features over 10,000 manually verified instructions and spanning diverse visual conditions, object types, and scales. We evaluate several state-of-the-art VLMs to assess performance on geospatial-specific challenges. The results indicate that although existing VLMs demonstrate potential, they face challenges when dealing with geospatial-specific tasks, highlighting the room for further improvements.
Notably, the best-performing LLaVa-OneVision achieves only 41.7% accuracy on MCQs, slightly more than GPT-4o, which is approximately double the random guess performance.

πŸ† Contributions

  • GEOBench-VLM Benchmark. We introduce GEOBench-VLM, a benchmark suite designed specifically for evaluating VLMs on geospatial tasks, addressing geospatial data challenges. It covers 8 broad categories and 31 sub-tasks with over 10,000 manually verified instructions.
  • Evaluation of VLMs. We provide a detailed evaluation of 13 state-of-the-art VLMs, including generic (open and closed-source) and geospatial-specific VLMs, highlighting their capabilities and limitations in geospatial analysis.
  • Analysis of Geospatial Task Performance. We analyze performance across a range of tasks, including scene classification, counting, change detection, relationship prediction, referring expression detection, segmentation, image captioning, disaster detection, and temporal analysis, among others, providing key insights that can help in improving VLMs for geospatial applications.

πŸ—‚οΈ Benchmarks Comparison

Dataset Comparison table

Table: Overview of Generic and Geospatial-specific Datasets & Benchmarks, detailing modalities (O=Optical, PAN=Panchromatic, MS=Multi-spectral, IR=Infrared, SAR=Synthetic Aperture Radar, V=Video, MI=Multi-image, BT=Bi-Temporal, MT=Multi-temporal), data sources (DRSD=Diverse RS Datasets, OSM=OpenStreetMap, GE=Google Earth, answer types (MCQ=Multiple Choice, SC=Single Choice, FF=Free-Form, BBox=Bounding Box, Seg=Segmentation Mask), and annotation types (A=Automatic, M=Manual).


πŸ” Dataset Annotation Pipeline

Our pipeline integrates diverse datasets, automated tools, and manual annotation. Tasks such as scene understanding, object classification, and non-optical analysis are based on classification datasets, while GPT-4o generates unique MCQs with five options: one correct answer, one semantically similar ``closest" option, and three plausible alternatives. Spatial relationship tasks rely on manually annotated object pair relationships, ensuring consistency through cross-verification. Caption generation leverages GPT-4o, combining image, object details, and spatial interactions with manual refinement for high precision.


πŸ“Š Results

Performance summary of VLMs. LLaVA-OneVision achieves the average accuracy (41.7%), slightly outperforming GPT-4o, which is relatively better in building counting, and general aircraft counting. EarthDial demonstrates strong results in scene classification. The overall results highlight VLMs' varying strengths across geospatial tasks, with even the best models achieving accuracy only slightly above double the random guess

Results Heatmap

Temporal Understanding Results

VLM performance on temporal geospatial tasks. Evaluation spans crop type classification, disaster type classification, farm pond change detection (CD), land use classification, and damaged building counting. EarthDial performs best in land use classification, while GPT-4o achieves better performance in disaster classification and damaged building counting. Qwen2-VL stands second in disaster classification.

ModelCrop Type ClassificationDisaster Type ClassificationFarm Pond Change DetectionLand Use ClassificationDamaged Building Count
EarthDial0.21820.57270.21050.66230.4362
GPT-4o0.18180.63000.17110.65250.5667
LLaVA-OneVision0.14550.45370.18420.58690.4810
Qwen2-VL0.10910.59910.19740.59670.5000

Reffering Expression Detection

Referring expression detection. We report Precision on 0.5 IoU and 0.25 IoU

ModelPrecision@0.5 IoUPrecision@0.25 IoU
Sphinx0.34080.5289
EarthDial0.24290.4139
GeoChat0.11510.2100
Ferret0.09430.2003
Qwen2-VL0.15180.2524
GPT-4o0.00870.0386
LHRS-Nova0.09300.2423
SkySenseGPT0.10820.3224

πŸ€– Qualitative Results

Scene Understanding: This illustrates model performance on geospatial scene understanding tasks, highlighting successes in clear contexts and challenges in ambiguous scenes. The results emphasize the importance of contextual reasoning and addressing overlapping visual cues for accurate classification.

Scene Understanding

Counting: The figure showcases model performance on counting tasks, where Qwen 2-VL, GPT-4o and LLaVA-One have better performance in identifying objects. Other models, such as Ferret, struggled with overestimation, highlighting challenges in object differentiation and spatial reasoning.

Counting

Object Classification: The figure highlights model performance on object classification, showing success with familiar objects like the "atago-class destroyer" and "small civil transport/utility" aircraft. However, models struggled with rarer objects like the ``murasame-class destroyer" and ``garibaldi aircraft carrier" indicating a need for improvement on less common classes and fine-grained recognition.

Object Classification

Event Detection: Model performance on disaster assessment tasks, with success in scenarios like 'fire' and 'flooding' but challenges in ambiguous cases like 'tsunami' and 'seismic activity'. Misclassifications highlight limitations in contextual reasoning and insufficient exposure on overlapping disaster features.

Event Detection

Spatial Relations: The figure demonstrates model performance on spatial relationship tasks, with success in close-object scenarios and struggles in cluttered environments with distant objects.

Spatial Relations


πŸ“œ Citation

If you find our work and this repository useful, please consider giving our repo a star and citing our paper as follows:

@inproceedings{danish2025geobenchvlm,
  title     = {GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks},
  author    = {Muhammad Sohail Danish and Muhammad Akhtar Munir and Syed Roshaan Ali Shah and Kartik Kuckreja and Fahad Shahbaz Khan and Paolo Fraccaro and Alexandre Lacoste and Salman Khan},
  booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
  year      = {2025},
  url       = {https://arxiv.org/abs/2411.19325},
  archivePrefix = {arXiv},
  eprint    = {2411.19325},
  primaryClass = {cs.CV}
}

πŸ“¨ Contact

If you have any questions, please create an issue on this repository or contact at muhammad.sohail@mbzuai.ac.ae.