SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
December 1, 2025 Β· View on GitHub
School of Artificial Intelligence, OPtics, and ElectroNics (iOPEN), Northwestern Polytechnical University
This is the official repository for paper "SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model". [paper] [SkyEye-968k]
π Accepted by ISPRS Journal of Photogrammetry and Remote Sensing π
Please share a STAR β if this project does help
You can focus on remote sensing multimodal large language model (Vision-Language) here
You can focus on multimodal large language model (Vision-Language) for UAV here
π’ Latest Updates
This is an ongoing project. We will be working on improving it.
- π¦ Chatbot, codebase, and model inference tutorial coming soon! π
- May-13-2025: SkyEyeGPT model checkpoint is released. [huggingface] π₯π₯ οΌThe Model Weight can be run directly with MiniGPT-v2οΌ
- Jan-19-2025: SkyEyeGPT paper is accepted by ISPRS. [paper] π₯π₯
- Jun-12-2024: RS instruction dataset SkyEye-968k is released. [huggingface] π₯π₯
- Jan-18-2024: paper is released. π₯π₯
- Jan-17-2024: A curated list about remote sensing multimodal large language model (Vision-Language) is created. π₯π₯
π¬ SkyEyeGPT: Remote Sensing Multi-modal Chatbot
The online demo will be released.
π Inference
We release the model weight in [huggingface]! The Model Weight can be run directly with MiniGPT-v2
SkyEyeGPT: Architecture
The model and checkpoint are coming soon! π
π SkyEye-968k: Unified RS Vision-Language Instruction
The download link of the unified remote sensing vision-language instruction dataset is here! π
Download link: https://huggingface.co/datasets/ZhanYang-nwpu/SkyEye-968k
π¦ Performance
ποΈ Visualization
1. Detailed description
2. Some testing samples of captioning, grounding, and VQA
ποΈ Qualitative results
1. Remote Sensing Visual Grounding
2. Remote Sensing Phrase Grounding
3. Remote Sensing Image Captioning
4. UAV Aerial Video Captioning
5. Remote Sensing Visual Question Answering
6. Remote Sensing Referring Expression Generation
7. Remote Sensing Scene Classification
π Quantitative results
1. Remote Sensing Image Captioning
2. UAV Aerial Video Captioning
3. Remote Sensing Visual Grounding
4. Remote Sensing Visual Question Answering
π Citation
@ARTICLE{zhan2025skyeyegpt,
title={SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model},
author={Yang Zhan and Zhitong Xiong and Yuan Yuan},
year={2025},
journal={ISPRS Journal of Photogrammetry and Remote Sensing},
volume = {221},
pages = {64-77}
}
π Acknowledgement
Our code is based on MiniGPT-4, shikra, and MiniGPT-v2. We sincerely appreciate their contributions and authors for releasing source codes. We are thankful to EVA and LLaMA2 for releasing their models as open-source contributions. I would like to thank Xiong zhitong and Yuan yuan for helping the manuscript. I also thank the School of Artificial Intelligence, OPtics, and ElectroNics (iOPEN), Northwestern Polytechnical University for supporting this work.
π€ Contact
If you have any questions about this project, please feel free to contact zhanyangnwpu@gmail.com.