SARChat

July 6, 2026 ยท View on GitHub

SARChat Logo

License HF Dataset HF Model ModelScope Dataset ModelScope Model ModelScope Open Weights arXiv

Introduction

SARChat-Bench-2M is the first large-scale multimodal dialogue dataset focusing on Synthetic Aperture Radar (SAR) imagery. It contains approximately 2 million high-quality SAR image-text pairs, supporting multiple tasks including scene classification, image captioning, visual question answering, and object localization. We conducted comprehensive evaluations on 16 state-of-the-art vision-language models (including Qwen2VL, InternVL2.5, and LLaVA), establishing the first multi-task benchmark in the SAR domain.

๐Ÿ“‘ Read more about SARChat in our paper.

๐Ÿš€ The SARChat model weights have been open-sourced and are available on ModelScope.

Overview & Model Performance

SARChat Tasks and Model Performance
Figure 1: Overview of SARChat's architecture (left) and comprehensive evaluation results showing model capabilities across different tasks (right)

Data Processing Workflow

SARChat Data Workflow
Figure 2: Data processing workflow of SARChat

Key Features

  • ๐ŸŒŸ 2M+ high-quality SAR image-text pairs
  • ๐Ÿ” Covers diverse scenes including marine, terrestrial and urban areas
  • ๐Ÿ“Š 6 task-specific benchmarks with fine-grained annotations
  • ๐Ÿค– Evaluated on 11 SOTA vision-language models
  • ๐Ÿ› ๏ธ Ready-to-use format with shape, count, location labels

Dataset Statistics

Tasks Statistics

Train Task Distribution Test Task Distribution
Figure 3: Distribution of tasks in training (left) and test (right) sets

TaskTrain SetTest Set
Classification81,78810,024
Fine-Grained Description46,1416,032
Instance Counting95,49311,704
Spatial Grounding94,45611,608
Cross-Modal Identification1,423,548175,565
Referring95,48611,703

Category Analysis

Train Categories Test Categories
Figure 4: Category distribution in training (left) and test (right) sets

Words Statistics

MetricValue
Total Words43,978,559
Total Sentences4,222,143
Average Caption Length10.66

Quick Start

๐Ÿค— Visit our Hugging Face dataset page for more details and examples.

๐Ÿ“ฆ Download the open-sourced SARChat weights from our ModelScope collection.

Results Showcase

SARChat Results
Figure 6: Example results from SARChat-InternVL2.5-8B model on various SAR vision-language tasks

The above figure demonstrates the capabilities of our SARChat-InternVL2.5-8B model across different tasks. The model shows strong performance in understanding complex SAR imagery, providing detailed descriptions, accurate counting, and precise spatial reasoning. These results highlight the model's ability to bridge the gap between SAR imagery and natural language understanding.

SARChat Models

We have trained and evaluated several models using the SARChat dataset:

OrganizationModelSizeLink
InternVLSARChat-InternVL2.51BLink
InternVLSARChat-InternVL2.52BLink
InternVLSARChat-InternVL2.54BLink
InternVLSARChat-InternVL2.58BLink
QwenVLSARChat-Qwen2VL2BLink
QwenVLSARChat-Qwen2VL7BLink
DeepSeekSARChat-DeepSeekVL1.3BLink
DeepSeekSARChat-DeepSeekVL7BLink
mPLUG-OwlSARChat-Owl31BLink
mPLUG-OwlSARChat-Owl32BLink
mPLUG-OwlSARChat-Owl37BLink
MicrosoftSARChat-Phi3V4.3BLink
Zhipu AISARChat-GLM-Edge2BLink
Zhipu AISARChat-GLM-Edge5BLink
LLaVA-TeamSARChat-LLaVA-1.57BLink
01.AISARChat-Yi-VL6BLink

Citation

If you use this dataset or our models in your research, please cite our paper.

@inproceedings{Ma2025SARChatBench2MAM,
  title={SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation},
  author={Zhiming Ma and Xiayang Xiao and Sihao Dong and Peidong Wang and HaiPeng Wang and Qingyun Pan},
  year={2025},
  url={https://api.semanticscholar.org/CorpusID:276287423}
}

Contact

For any questions or feedback, please contact:


If you find SARChat useful, please consider giving it a star โญ