README.md

September 9, 2026 · View on GitHub

Image

Awesome Remote Sensing Vision-Language Datasets

DREAMS@ECNU  VisionXLab@SJTU 

Awesome Website Issue's Welcome Paper App Demo Agent Skill


By providing curated, high-quality dataset recommendations, this repository addresses the critical data needs for developing advanced Large Vision-Language Models (LVLMs) in Remote Sensing.

Inspired by the JE project, we have adopted GitHub Issues to manage our datasets. With the help of Labels, it's easy to filter the datasets we need. I won't go into extensive detail about the benefits of using GitHub Issues as a database, but if you're an open-source enthusiast, trust me, you'll love it.

We encourage researchers to submit outstanding results that we may have missed to the project's Issues. Please provide the essential information based on the provided template. Subsequently, we will add appropriate labels and update the README page accordingly. We also encourage users to join the discussion under each issue—sharing their experiences, feedback, and whether they would recommend the dataset to others.

🥳 New

🔥🔥🔥 Last Updated on 2026.07.31 🔥🔥🔥

  • 2026.06.25: FUSAR-GEOVL-1M
  • 2026.06.25: RS-Neg
  • 2026.06.08: Sky-VT-300K
  • 2026.05.21: 🚀🚀🚀 Accepted by IEEE Geoscience and Remote Sensing Magazine (GRSM)
  • 2026.04.28: We release the GeoChef SKill for your OpenClaw/Hermes/Coze/WorkBuddy! Excellent, Highly Recommended!!
  • 2026.04.22: We release our Agent (test version) to assist researchers in the RS vision-language field with dataset-related problems.

dara_type

Usage Demo:

  • Human Annotation:
  • Generated by GPT-4V:
  • Ultra High Resolution:
  • Temporal Understanding:

Summary of Contents

Comprehensive Data

YearVenueCountryNameDownloadMore
2024CVPRUnited Arab EmiratesGeoChat-InstructStar
2024ISPRSChinaSkyEye-968kStar
2024TGRSChinaMMRS-1MStar
2024NeurIPSSaudi ArabiaVRSBench-TrainStar
2024arXivChinaRS-GPT4VStar
2024TGRSChinaRSVPStar
2024ECCVChinaLHRS-Align & LHRS-InstructStar
2024arXivChinaFIT-RSStar
2025CVPRUnited StatesEarthDial-InstructStar
2025RSChinaDDFAVStar
2025ICMLUnited Arab EmiratesGeoPixelDStar
2025arXivChinaFalcon_SFTStar
2025AAAIChinaVersaD & HnstD & VariousRS-InstructStar
2025ICLRUnited StatesTEOChatlasStar
2025ICASSPChinaChangeChat-87kStar
2025GRSMChinaGeoPixInstructStar
2025ICLRJapanSARLANG-1MStar
2025arXivChinaM-RSVPStar
2025arXivChinaRS-VL3MNaN
2025ISPRSChinaLHRS-Align-Recap & LHRS-Instruct-PlusStar
2025arXivChinaFUSAR-GEOVL-1MStar
2025NeurIPSChinaScoreRSStar
2025NeurIPSJapanDisasterM3Star
2025arXivChinaSARVLM-1MNAN
2025arXivChinaRS-EoT-4KStar
2026arXivUnited Arab EmiratesOpenEarthAgent DatasetStar
2026arXivGermanyBigEarthNet.txtNAN

Comprehensive Benchmarks

YearVenueCountryNameDownloadMore
2024ICLRUnited StatesNAIP-OSMlink
2024CVPRUnited Arab EmiratesGeoChat-BenchStar
2024ECCVChinaLHRS-BenchStar
2024NeurIPSSaudi ArabiaVRSBenchStar
2024arXivChinaFIT-RSFGStar
2024arXivChinaFIT-RSRCStar
2024arXivUnited StatesVLEO-BenchStar
2025ICCVUnited Arab EmiratesGEOBench-VLMStar
2025AAAIChinaUrBenchStar
2025CVPRChinaXLRS-BenchStar
2025NeurIPSChinaCHOICEStar
2025ISPRSChinaRSIEvalStar
2025TGRSChinaAirSpatialStar
2025arXivChinaSARChat-Bench-2MStar
2025arXivChinaREOBenchStar
2025arXivUnited Arab EmiratesThinkGeoStar
2025arXivChinaA2SeekStar
2026arXivChinaGeoReason-BenchStar
2026arXivGermanyBigEarthNet.txtNAN
2026ECCVChinaRS-NegStar

Task-specific Data

Image Captioning / Retrieval

Image Captioning and Retrieval tasks share the same dataset, consisting of images and their corresponding descriptions. The key difference is that captioning allows multiple textual descriptions per image.

YearVenueCountryNameDownloadMore
2016CITSChinaUCM-Captionslink
2016CITSChinaSydney-Captionslink
2017TGRSChinaRSICDStar
2021TGRSChinaRSITMDStar
2022TGRSChinaNWPU-CaptionsStar
2022TGRSItalyUAV-CaptionsN/A
2022TGRSChinaLEVIR-CCStar
2022RSSaudi ArabiaCapERAStar
2023ISPRSChinaRSICapStar
2023ICCVEuropean UnionLAION-EOlink
2024TGRSChinaRemoteCLIPStar
2024TGRSChinaRS5MStar
2024AAAIUnited StatesSkyScriptStar
2025ISPRSChinaRSTellerStar
2024NeurIPSChinaMMM-RSStar
2024droneChinaMOCOStar
2025ISPRSChinaLuoJiaHOGN/A
2025TGRSChinaFITN/A
2025arXivChinaLRS2MStar
2025arXivGermanyGeoLangBind-2MStar
2025GRSMChinaGit-10MStar
2025arXivGreeceGAIAStar
2025ICCVChinaCVG-TextStar
2025PKDDGermanyLlama3-SSL4EOS12Star
2025arXivChinaSPIEStar
2025arXivChinaGLEAMStar
2025GRSMChinaStreet2Sat-TextN/A
2025GRSMChinaCVACT-TextN/A
2025arXivChinaSAR-TEXTStar
2025NeurIPSChinaRSCCStar
2025AAAIAustraliaLandsat30-AU-CAPStar
2025ISPRSChinaSARCAPStar
2025TGRSChinaMMSARNaN
2026TPAMIChinaSkyFindStar

Visual Question Answering (VQA)

YearVenueCountryNameDownloadMore
2020TGRSNetherlandsRSVQA-LR & RSVQA-HRlink
2021IGARSSFranceRSVQAxBENlink
2021AccessUnited StatesFloodNetStar
2021TGRSChinaRSIVQAStar
2022TGRSGermanyCDVQAStar
2022IJRSSaudi ArabiaVQA-TextRSN/A
2022TGRSChinaCRSVQAStar
2024TGRSChinaRemoteCountStar
2024AAAIChinaEarthVQAStar
2024arXivUnited Arab EmiratesGeoLLaVAStar
2024arXivChinaQAG-360KStar
2024ISPRSNetherlandsHRVQAlink
2024ISPRSChinaOSVQAlink
2025ICLRChinaMME-RealWorldStar
2025arXivChinaLRS-VQAStar
2025arXivChinaAgroMindStar
2025ISPRSChinaAVI-MathStar
2025arXivChinaEarth-BenchStar
2025NeurIPSChinaVICoT-HRSCNaN
2025arXivChinaGeoPlan-BenchStar
2025arXivChinaLRS-GROStar
2025AAAIAustraliaLandsat30-AU-VQAStar
2025ISPRSJapanCitySetNaN
2025ICMLUnited StatesUnivEARTHNaN
2025arXivJapanGeo3DVQAStar
2025arXivChinaUrbanVideo-BenchStar
2025arXivChinaBEDIStar
2025arXivChinaKnowFlow-BenchNaN
2025arXivChinaMM-UAVBenchStar
2025arXivChinaLRS-VQA-ZoomStar
2026arXivChinaUHR-CoZStar
2026CVPRJapanGeoMMBenchStar
2026ArXivChinaFUSAR-GEOVL-1MNaN

Visual Grounding

YearVenueCountryNameDownloadMore
2022MMChinaRSVGlink
2023TGRSChinaDIOR-RSVGStar
2024ECCVSingaporeGeoText-1652Star
2024TGRSChinaRSVG-HRStar
2024TGRSChinaOPT-RSVGStar
2024TGRSGermanyRRSIS (RefSegRS)link
2024CVPRChinaRRSIS-DStar
2024arXivSingaporeAAVGStar
2024arXivSingaporerefGeoStar
2024arXivChinaRISBenchStar
2025ICCVChinaAerialVGStar
2025arXivChinaEarthReasonStar
2025arXivUnited StatesGRESStar
2025arXivChinaRefDroneStar
2025MMChinaRemoteSAM-270KStar
2025arXivChinaNWPU-ReferStar
2025arXivChinaGeoSeg-1MStar
2025AAAIChinaRIS-LADStar
2025arXivChinaGRASP-1KNaN
2025arXivChinaLaSeRSStar
2026CVPRChinaSky-VT-300KStar

Meta Data

All remote sensing datasets used to construct the above-mentioned VL datasets.

Note that we do not include the metadata in the issue.

If you want to see how many datasets used the xBD metadata during construction,

Usage Demo: https://github.com/zytx121/Awesome-RS-SFT-Data/issues?q=is%3Aissue%20state%3Aopen%20xBD

Classification

YearVenueKeywordsNameDownload
2010GISUCM (UCMerced)link
2011TGRSWHU-RS19link
2015GRSLRSSCN7Star
2015TGRSSIRI-WHUlink
2015SIGSPATIALSAT-4 & SAT-6link
2015TGRSSydneyN/A
2016JARSRS C11link
2017TGRSRSD46-WHUlink
2017TGRSAIDlink
2017Proc. IEEENWPU-RESISC45link
2018WebsiteHefeilink
2018DPHurricane Damagelink
2018TGRSOPTIMAL31link
2018ISPRSPatternNetlink
2018CVPRfMoWStar
2019JSTARSEuroSATlink
2019IGARSSBigEarthNetlink
2019GRSMSo2Satlink
2020WebsiteAiRoundlink
2020Websiteairplane_detlink
2020SensorCLRSlink
2020SensorRSI-CBStar
2020ISPRSMLRSNetStar
2021Websiteship_detlink
2021JSTARSMillionAIDlink
2021TGRSMultiSceneStar
2021JSTARSNaSC-TG2link
2021KSEDSCRStar
2021RSFGSCR-42Star
2021IJAEOGWHU-OPT-SARStar
2022arXivMETER-MLlink
2022RSMRSSC2.0link
2023ICCVWSATINlink
2023DICTAFireRiskStar
2025JASMEETStar

Detection

  • SAR: Synthetic Aperture Radar
  • IR: Infrared
  • Attrs: Object-level Attribute Understanding
YearVenueKeywordsNameDownload
2012TPAMISZTAKIlink
2014ISPRSNWPU VHR-10link
2015ICIPUCAS-AODStar
2016ECCVS-Dronelink
2016ECCVCOWClink
2017ICPRAMHRSC2016link
2017TGRSRSODStar
2017TIPLEVIRlink
2017ICCVCARPKlink
2017BIGSARDATASARSSDDlink
2018CVPRDOTAlink
2018KaggleASDlink
2018arXivxViewlink
2018WebsiteDeepGlobe Detectionlink
2018ICIPITCVDlink
2019TGRSHRRSDStar
2019JRSARAIR-SARShip-2.0link
2019MEEAerialAnimallink
2019WebsiteSPCDlink
2020ISPRSDIORlink
2020ICRAAU-AIRlink
2020AccessSARHRSIDStar
2020WebsiteOceanic-Shiplink
2020WebsiteRarePlaneslink
2021WebsiteForest Damageslink
2021RSS2-SHIPSlink
2021JSTARSShipRSImagerNetlink
2021TPAMIVisDroneStar
2021WebsiteSea-Shippinglink
2021WebsiteIRInfrared-Securitylink
2021WebsiteAerial-Mancarlink
2021WebsiteDouble-Light-Vehiclelink
2021WebsiteMarine Debrislink
2021PCBINEON-Treelink
2021RSSARSRSDDStar
2022JRSARMSARlink
2022JRSMAR20link
2022ISPRSFAIR1Mlink
2022IJGIVHRShipslink
2022TGRSLEVIR-ShipStar
2023TPAMISODA-Alink
2023SDHIT-UAVStar
2024WebsiteDeforestationlink
2024TPAMIGLH-BridgeStar
2024NeurIPSSARDet-100KStar
2025TPAMISGGSTARStar
2025CVPRSARRSARStar
2025arXivAttrsEVAttrs-95KStar

Segmentation

YearVenueKeywordsNameDownload
2017IGARSSInrialink
2018RSDLRSDlink
2018WebsiteVaihingenlink
2018WebsitePotsdamlink
2018WebsiteTorontolink
2018WebsiteDeepGlobe Land Coverlink
2019CVPRWiSAIDlink
2019ICCVSkyScapeslink
2020FAICrowdAIlink
2020WebsiteBHP Watertankslink
2020RSEGID-15link
2020arXivHi-UCDlink
2020ISPRSUAVidlink
2021TGRSGEONRWlink
2021arXivLoveDAStar
2022MLMiniFrancelink
2023JRSGlobe230klink
2023NeurIPSSAMRSStar
2023SRSTARCOPStar
2024arXivChatEarthNetStar
2024arXivFineGripN/A

Change Detection

YearVenueKeywordsNameDownload
2018ArchivesCDDlink
2018TGRSWHU-CDlink
2019arXivxBDStar
2019CVIUHRSCDlink
2020ISPRSDSIFNStar
2020RSLEVIR-CDlink
2020arXivSECONDlink
2021ICCVPASTISStar
2021RSLEVIR-CD+Star
2021RSS2LookingStar
2021TGRSSYSU-CDStar
2021CVPRWQFabriclink
2022JSTARSMSBCStar
2022JSTARSMSOSCDStar
2022ISPRSNJDSlink
2023GRSLEGY BCDStar
2023arXivFPCDlink
2023ISPRSGVLMStar
2024JRSMtSCCDlink
2024ECCVMUDSlink
2024TGRSLEVIR-MCIStar

Others

  • SGG: Scene Graph Generation
  • Geo-Loc.: Geo-Localization
  • ER: Event Recognition
YearVenueKeywordsNameDownload
2017CVPRGeo-Loc.CVUSAStar
2018arXivSEN1-2link
2019CVPRGeo-Loc.CVACTStar
2020CVPRWSpaceNet6link
2020MMGeo-Loc.University-1652Star
2020GRSMERERAStar
2023GRSMDFC2023link
2023ICCVSatlasPretrainlink
2023ICCVGeoPileStar
2023GRSMSSL4EO-S12Star
2024ISCRAMQuakeSetlink
2024LNCSTreeSatAI-TSStar
2025arXivOpenLandMapStar
2025arXivOpenEarthMap-SARlink
2025ICCVCrossText2LocStar

Papers

Multimodal Large Language Models

YearVenueKeywordsNameDownload
2024CVPRGeoChat: Grounded Large Vision-Language Model for Remote SensingStar
2024ISPRSSkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language ModelStar
2024TGRSEarthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domainStar
2024ECCVLHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language ModelStar
2024arXivPopeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing ImageryN/A
2024arXivLarge Language Models for Captioning and Retrieving Remote Sensing ImagesN/A
2024RSRS-LLaVA: A Large Vision-Language Model for Joint Captioning and Question Answering in Remote Sensing ImageryStar
2024arXivSkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingStar
2024TGRSEarthMarker: A Visual Prompt Learning Framework for Region-level and Point-level Remote Sensing Imagery ComprehensionStar
2024arXivAquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image ComprehensionN/A
2024arXivLarge Vision-Language Models for Remote Sensing Visual Question AnsweringN/A
2024arXivGeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual GroundingStar
2025ISPRSLHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language InterpretationStar
2024arXivGeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote SensingN/A
2024TGRSRingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and grounded tasksN/A
2024arXivRSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of ExpertsStar
2024arXivUniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language ModelsStar
2024arXivREO-VLM: Transforming VLM to Meet Regression Challenges in Earth ObservationN/A
2025TGRSRS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question AnsweringN/A
2025ICLRTEOChat: A Large Vision-Language Assistant for Temporal Earth Observation DataStar
2025AAAIVHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisStar
2025CVPREarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesStar
2025GRSMGeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote SensingStar
2025ICMLGeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingStar
2025arXivQuality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsN/A
2025arXivGeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation ModelsStar
2025IGARSSLMMRotate: A Simple Aerial Detection Baseline of Multimodal Language ModelsStar
2025arXivWhen Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token PruningStar
2025arXivFalcon: A Remote Sensing Vision-Language Foundation ModelStar
2025arXivOmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial IntelligenceN/A
2025CVPRXLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?Star
2025arXivEagleVision: Object-level Attribute Multimodal LLM for Remote SensingStar
2025arXivSegEarth-R1: Geospatial Pixel Reasoning via Large Language ModelStar
2025arXivEarthGPT-X: Enabling MLLMs to Flexibly and Comprehensively Understand Multi-Source Remote Sensing ImageryNaN
2025ISPRSRsgpt: A remote sensing vision language model and benchmarkStar
2025arXivSARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image InterpretationStar
2025arXivLISAt: Language-Instructed Segmentation Assistant for Satellite ImageryStar
2025GRSMImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAGStar
2025arXivRingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal ReasoningNaN
2025arXivVectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMsNaN
2026AAAIRemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowNaN
2025arXivRS-OOD: A Vision-Language Augmented Framework for  Out-of-Distribution Detection in Remote SensingNaN
2025arXivMSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial ReasoningNaN
2025ECML PKDD 2025Beyond the Visible: Multispectral Vision-Language Learning for Earth ObservationStar
2025IEEESatellite Image Synthesis From Street View With Fine-Grained Spatial Textual Guidance: A novel frameworkNaN
2025arXivRS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelNaN
2025arXivGeo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-TuningStar
2025arXivLook Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom SearchStar
2025arXivGeoZero: Incentivizing Reasoning from Scratch on Geospatial ScenesNaN

Vision-Language Pre-training Models

YearVenueKeywordsNameDownload
2023TGRSParameter-Efficient Transfer Learning for Remote Sensing Image–Text RetrievalStar
2023JAGRS-CLIP: Zero Shot Remote Sensing Scene Classification via Contrastive Vision-Language SupervisionStar
2024TGRSRS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Vision-Language Foundation Model for Remote SensingStar
2024ICLRRemote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentProject
2024AAAISkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingStar
2024arXivMind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal AlignmentN/A
2024TGRSRemoteCLIP: A Vision Language Foundation Model for Remote SensingStar
2025arXivLRSCLIP: A Vision-Language Foundation Model for Aligning Remote Sensing Image with Longer TextStar
2025PKDDBeyond the Visible: Multispectral Vision-Language Learning for Earth ObservationStar
2025ICCVCopernicus: Towards a Unified Copernicus Foundation Model for Earth VisionStar
2025arXivSkySense V2: A Unified Foundation Model for Multi-modal Remote SensingN/A
2025arXivCGEarthEye: A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite ConstellationN/A

Intelligent Agents

YearVenueKeywordsNameDownload
2023arXivTree-GPT: Modular Large Language Model Expert System for Forest Remote Sensing Image Understanding and Interactive AnalysisN/A
2024IGARSSRemote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual ModelsStar
2024ICLRWEvaluating Tool-Augmented Agents in Remote Sensing PlatformsN/A
2024arXivGeoLLM-Engine: A Realistic Environment for Building Geospatial CopilotsN/A
2024arXivRS-Agent: Automating Remote Sensing Tasks through Intelligent AgentsN/A
2024TGRSChange-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and AnalysisStar
2025CVPRPEACE: Empowering Geologic Map Holistic Understanding with MLLMsStar
2025TGRSAirSpatialBot: A Spatially Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognition and RetrievalStar

Survey

YearVenueKeywordsNameDownload
2023IGARSSAn Agenda for Multimodal Foundation Models for Earth ObservationN/A
2023GISWULarge Remote Sensing Model: Progress and ProspectsN/A
2023JSTARSBrain-Inspired Remote Sensing Foundation Models and Open Problems: A Comprehensive SurveyN/A
2023arXivOn the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning ApplicationsN/A
2024GRSMVision-Language Models in Remote Sensing: Current Progress and Future TrendsN/A
2024arXivOn the Foundations of Earth and Climate Foundation ModelsN/A
2024arXivTowards Vision-Language Geo-Foundation Model: A SurveyStar
2024SBFoundation model for generalist remote sensing intelligence: Potentials and prospectsN/A
2024arXivAdvancements in Visual Language Models for Remote Sensing: Datasets, Capabilities, and Enhancement TechniquesStar
2024arXivFoundation Models for Remote Sensing and Earth Observation: A SurveyStar
2024GRSMWhen Geoscience Meets Foundation Models: Toward a general geoscience artificial intelligence systemN/A
2025JAGTowards the next generation of Geospatial Artificial IntelligenceN/A
2025InnovationUnleashing the potential of remote sensing foundation models via bridging data and computility islandsN/A
2025arXivA Survey on Remote Sensing Foundation Models: From Vision to MultimodalityStar
2025IEEERegression in Earth Observation: Are vision–language models up to the challenge?N/A

Exploration

YearVenueKeywordsNameDownload
2023arXivGPT4GEO: How a Language Model Sees the World's GeographyStar
2023arXivCharting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsStar
2023arXivThe Potential of Visual ChatGPT For Remote SensingN/A

Awesome Lists

NameDownload
satellite-image-deep-learningStar
Awesome Satellite Imagery DatasetsStar
Awesome-Remote-Sensing-Foundation-ModelsStar
Awesome-Remote-Sensing-Multimodal-Large-Language-ModelsStar
Awesome Visual Language Models papers and resources for Earth ObservationStar
Awesome remote sensing vision language modelsStar
Vision-Language Geo-Foundation ModelsStar
Awesome Remote Sensing Vision-Language Models & PapersStar
awesome-remote-image-captioningStar
Foundation Models for Remote Sensing and Earth ObservationStar
Awesome VLMs in RSStar
Remote Sensing Foundation ModelsStar

Citation

If you find our survey and repository useful for your research project, please consider citing our paper:

@ARTICLE{zhou2026geochef,
  author={Zhou, Yue and Zhao, Shujun and Yang, Xue and Li, Ruigang and Zhang, Tianwen and Lan, Mengcheng and Chen, Chaofeng and Ma, Lingfei and He, Hongjie and Li, Jonathan},
  journal={IEEE Geoscience and Remote Sensing Magazine}, 
  title={Data-Driven Vision-Language Models for Remote Sensing: A survey}, 
  year={2026},
  volume={},
  number={},
  pages={2-37},
  keywords={Modeling;Visual systems;Cognition;Cognitive systems;Remote sensing;Tuning;Grounding;Visualization;Large language models;Vision language model},
  doi={10.1109/MGRS.2026.3696441}}

Contact

yzhou@geoai.ecnu.edu.cn