Awesome Remote Sensing Foundation Models

May 7, 2026 · View on GitHub

Maintenance Awesome GitHub watchers GitHub stars GitHub forks

Awesome Remote Sensing Foundation Models

:star2:A collection of papers, datasets, benchmarks, code, and pre-trained weights for Remote Sensing Foundation Models (RSFMs).

📢 Latest Updates

:fire::fire::fire: Last Updated on 2026.05.06 :fire::fire::fire:

Table of Contents

Remote Sensing Vision Foundation Models

AbbreviationTitlePublicationPaperCode & Weights
GeoKRGeographical Knowledge-Driven Representation Learning for Remote Sensing ImagesTGRS2021GeoKRlink
-Self-Supervised Learning of Remote Sensing Scene Representations Using Contrastive Multiview CodingCVPRW2021Paperlink
GASSLGeography-Aware Self-Supervised LearningICCV2021GASSLlink
SeCoSeasonal Contrast: Unsupervised Pre-Training From Uncurated Remote Sensing DataICCV2021SeColink
RSPAn Empirical Study of Remote Sensing PretrainingTGRS2022RSPlink
MATTERSelf-Supervised Material and Texture Representation Learning for Remote Sensing TasksCVPR2022MATTERlink
-Self-supervised Vision Transformers for Land-cover Segmentation and ClassificationCVPRW2022Paperlink
DINO-MMSelf-Supervised Vision Transformers for Joint SAR-Optical Representation LearningIGARSS2022DINO-MMlink
GeCoGeographical Supervision Correction for Remote Sensing Representation LearningTGRS2022GeColink
RingMoRingMo: A remote sensing foundation model with masked image modelingTGRS2022RingMoCode
RS-BYOLSelf-Supervised Learning for Invariant Representations From Multi-Spectral and SAR ImagesJSTARS2022RS-BYOLlink
CSPTConsecutive Pre-Training: A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing DomainRS2022CSPTlink
RVSAAdvancing plain vision transformer toward remote sensing foundation modelTGRS2022RVSAlink
SatMAESatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryNeurIPS2022SatMAElink
DINO-MCExtending Global-Local View Alignment for Self-Supervised Learning with Remote Sensing ImageryArxiv2023DINO-MClink
CMIDCMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image UnderstandingTGRS2023CMIDlink
PrestoLightweight, Pre-trained Transformers for Remote Sensing TimeseriesArxiv2023Prestolink
ASTAST: Adaptive Self-supervised Transformer for Optical Remote Sensing RepresentationISPRS JPRS2023ASTnull
TOVTOV: The original vision model for optical remote sensing image understanding via self-supervised learningJSTARS2023TOVlink
CACoChange-Aware Sampling and Contrastive Learning for Satellite ImagesCVPR2023CAColink
IaI-SimCLRMulti-Modal Multi-Objective Contrastive Learning for Sentinel-1/2 ImageryCVPRW2023IaI-SimCLRnull
-A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionIGARSS2023Papernull
SatLasSatlasPretrain: A Large-Scale Dataset for Remote Sensing Image UnderstandingICCV2023SatLaslink
GFMTowards Geospatial Foundation Models via Continual PretrainingICCV2023GFMlink
Scale-MAEScale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningICCV2023Scale-MAElink
PrithviFoundation Models for Generalist Geospatial Artificial IntelligenceArxiv2023Prithvilink
RingMo-SenseRingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution DisentanglingTGRS2023RingMo-Sensenull
EarthPTEarthPT: a time series foundation model for Earth ObservationNeurIPS2023 CCAI workshopEarthPTlink
CROMACROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersNeurIPS2023CROMAlink
Cross-Scale MAECross-Scale MAE: A Tale of Multiscale Exploitation in Remote SensingNeurIPS2023Cross-Scale MAElink
USatUSat: A Unified Self-Supervised Encoder for Multi-Sensor Satellite ImageryArxiv2023USatlink
AIEarthAnalytical Insight of Earth: A Cloud-Platform of Intelligent Computing for Geospatial Big DataArxiv2023AIEarthlink
GeRSPGeneric Knowledge Boosted Pretraining for Remote Sensing ImagesTGRS2024GeRSPGeRSP
SMLFRGenerative ConvNet Foundation Model With Sparse Modeling and Low-Frequency Reconstruction for Remote Sensing Image InterpretationTGRS2024SMLFRlink
RingMo-liteRingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid FrameworkIEEE TGRS2024RingMo-litenull
U-BARNSelf-Supervised Spatio-Temporal Representation Learning of Satellite Image Time SeriesJSTARS2024Paperlink
SpectralGPTSpectralGPT: Spectral Remote Sensing Foundation ModelTPAMI2024SpectralGPTlink
SwiMDiffSwiMDiff: Scene-Wide Matching Contrastive Learning With Diffusion Constraint for Remote Sensing ImageTGRS2024SwiMDiffnull
DOFANeural Plasticity-Inspired Multimodal Foundation Model for Earth ObservationArxiv2024DOFAlink
-Masked Feature Modeling for Generative Self-Supervised Representation Learning of High-Resolution Remote Sensing ImagesIEEE JSTARS2024Papernull
BFMA Billion-scale Foundation Model for Remote Sensing ImagesIEEE JSTARS2024BFMnull
ClayClay Foundation ModelArxiv2024nulllink
HydroHydro--A Foundation Model for Water in Satellite ImageryArxiv2024nulllink
S2MAES2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataCVPR2024S2MAEnull
SatMAE++Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryCVPR2024SatMAE++link
msGFMBridging Remote Sensors with Multisensor Geospatial Foundation ModelsCVPR2024msGFMlink
SkySenseSkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryCVPR2024SkySenselink
MTPMTP: Advancing Remote Sensing Foundation Model via Multi-Task PretrainingIEEE JSTARS2024MTPlink
RS-DFMRS-DFM: A Remote Sensing Distributed Foundation Model for Diverse Downstream TasksArxiv2024RS-DFMnull
OFA-NetOne for All: Toward Unified Foundation Models for Earth VisionIGARSS2024OFA-Netnull
-Lightweight and Efficient: A Family of Multimodal Earth Observation Foundation ModelsIGARSS2024Papernull
MM-VSFTowards Knowledge Guided Pretraining Approaches for Multimodal Foundation Models: Applications in Remote SensingArxiv2024MM-VSFnull
LeMeViTLeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image InterpretationIJCAI2024LeMeViTlink
SAR-JEPAPredicting Gradient is Better: Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive ArchitectureISPRS JPRS2024SAR-JEPAlink
DeCURDeCUR: decoupling common & unique representations for multimodal self-supervisionECCV2024DeCURlink
MMEarthMMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation LearningECCV2024MMEarthlink
OmniSatOmniSat: Self-Supervised Modality Fusion for Earth ObservationECCV2024OmniSatlink
MA3EMasked Angle-Aware Autoencoder for Remote Sensing ImagesECCV2024MA3Elink
SoftConMulti-Label Guided Soft Contrastive Learning for Efficient Earth Observation PretrainingTGRS2024SoftConlink
PISPretrain a Remote Sensing Foundation Model by Promoting Intra-instance SimilarityTGRS2024PISlink
FG-MAEFeature Guided Masked Autoencoder for Self-Supervised Learning in Remote SensingIEEE JSTARS2024FG-MAElink
-A Multimodal Unified Representation Learning Framework With Masked Image Modeling for Remote Sensing ImagesIEEE TGRS2024Papernull
OReole-FMOReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagerySIGSPATIAL2024OReole-FMnull
SatVision-TOASatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing ImageryArxiv2024SatVision-TOAlink
Prithvi-EO-2.0Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation ApplicationsArxiv2024Prithvi-EO-2.0link
A2-MAEA2-MAE: A spatial-temporal-spectral unified remote sensing pre-training method based on anchor-aware masked autoencoderIEEE TGRS2025A2-MAEnull
FoMoFoMo: Multi-Modal, Multi-Scale and Multi-Task Remote Sensing Foundation Models for Forest MonitoringAAAI2025FoMolink
PIEViTPattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote SensingIEEE TGRS2025PIEViTnull
SARATR-XSARATR-X: Toward Building a Foundation Model for SAR Target RecognitionIEEE TIP2025SARATR-Xlink
SatMambaSatMamba: Development of Foundation Models for Remote Sensing Imagery Using State Space ModelsArxiv2025SatMambalink
DynamicVisDynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation ModelsArxiv2025DynamicVislink
SUMMITSUMMIT: A SAR foundation model with multiple auxiliary tasks enhanced intrinsic characteristicsIJAEO2025SUMMITnull
SenPa-MAESenPa-MAE: Sensor Parameter Aware Masked Autoencoder for Multi-Satellite Self-Supervised PretrainingLNCS2025SenPa-MAElink
HyperSIGMAHyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelIEEE TPAMI2025HyperSIGMAlink
HyperSLHyperSL: A Spectral Foundation Model for Hyperspectral Image InterpretationIEEE TGRS2025HyperSLlink
TiMoTiMo: Spatiotemporal Foundation Model for Satellite Image Time SeriesArxiv2025TiMolink
PanopticonPanopticon: Advancing Any-Sensor Foundation Models for Earth ObservationCVPRW2025 (EarthVision Best Paper)Panopticonlink
HyperFreeHyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryCVPR2025HyperFreelink
SpectralEarthSpectralEarth: Training Hyperspectral Foundation Models at ScaleIEEE JSTARS2025SpectralEarthlink
WV-NetWV-Net: A foundation model for SAR WV-mode satellite imagery trained using contrastive self-supervised learning on 10 million imagesAIES2025WV-Netnull
AnySatAnySat: One Earth Observation Model for Many Resolutions, Scales, and ModalitiesCVPR2025AnySatlink
TerraFMTerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationArxiv2025TerraFMlink
GalileoGalileo: Learning Global & Local Features of Many Remote Sensing ModalitiesICML2025 TerraBytes WorkshopGalileolink
DeepAndesDeepAndes: A Self-Supervised Vision Foundation Model for Multispectral Remote Sensing Imagery of the AndesIEEE JSTARS2025DeepAndeslink
RingMambaRingMamba: Remote Sensing Multisensor Pretraining With Visual State Space ModelIEEE TGRS2025RingMambalink
RingMo-AerialRingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive LearningIEEE TPAMI2025RingMo-Aerialnull
SeaMoSeaMo: A Multi-Seasonal and Multimodal Remote Sensing Foundation ModelInformation Fusion2025SeaMonull
MoSAiCMoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote SensingIEEE Sensors Journal 2025MoSAiCnull
CGEarthEyeCGEarthEye: A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite ConstellationArxiv2025CGEarthEyenull
AlphaEarthAlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label dataArxiv2025AlphaEarthlink
SkySense++A semantic-enhanced multi-modal remote sensing foundation model for Earth observationNature Machine Intelligence 2025SkySense++link
CtxMIMCtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image UnderstandingACM TOMM2025CtxMIMnull
ViTPVisual Instruction Pretraining for Domain-Specific Foundation ModelsArxiv2025ViTPlink
SatDiFuserCan Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?ICCV2025SatDiFuserlink
FedSenseTowards Privacy-preserved Pre-training of Remote Sensing Foundation Models with Federated Mutual-guidance LearningICCV2025FedSensenull
Copernicus-FMTowards a Unified Copernicus Foundation Model for Earth VisionICCV2025Copernicus-FMlink
TerraMindTerraMind: Large-Scale Generative Multimodality for Earth ObservationICCV2025TerraMindlink
SelectiveMAEHarnessing Massive Satellite Imagery with Efficient Masked Image ModelingICCV2025SelectiveMAElink
SMARTIESSMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing ImagesICCV2025SMARTIESlink
SkySense V2SkySense V2: A Unified Foundation Model for Multi-modal Remote SensingICCV2025SkySense V2null
RS-vHeatRS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelICCV2025RS-vHeatnull
RoMARoMA: Scaling up Mamba-based Foundation Models for Remote SensingNeurIPS2025RoMAlink
GeoLinkGeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap DataNeurIPS2025GeoLinklink
CrossEarthCrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic SegmentationIEEE TPAMI2025CrossEarthlink
PhySwinPhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth ObservationNeurIPS2025PhySwinnull
FlexiMoFlexiMo: A Flexible Remote Sensing Foundation ModelIEEE TGRS2026FlexiMonull
MAPEXMAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation ModelsIEEE TGRS2026MAPEXlink
-A Complex-Valued SAR Foundation Model Based on Physically Inspired Representation LearningIEEE TIP2026Papernull
MAESTROMAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation DataWACV2026MAESTROlink
RingMoERingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image InterpretationIEEE TPAMI2026RingMoEnull
AllianceAlliance: All-in-One Spectral-Spatial-Frequency Awareness Foundation ModelIEEE TPAMI2026Alliancenull
THORTHOR: A Versatile Foundation Model for Earth Observation Climate and Society ApplicationsArxiv2026THORlink
AgriFMAgriFM: A multi-source temporal remote sensing foundation model for Agriculture mappingRSE2026Paperlink
SIGMAESIGMAE: A Spectral-Index-Guided Foundation Model for Multispectral Remote SensingArxiv2026SIGMAElink
CrossEarth-SARCrossEarth-SAR: A SAR-Centric and Billion-Scale Geospatial Foundation Model for Domain Generalizable Semantic SegmentationArxiv2026CrossEarth-SARlink
NeighborMAENeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders PretrainingCVPR2026NeighborMAEnull
MOMOMOMO: Mars Orbital Model Foundation Model for Mars Orbital ApplicationsCVPR2026MOMOlink
TESSERATESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisCVPR2026TESSERAlink
OlmoEarthOlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationCVPR2026OlmoEarthlink
RAMENRAMEN: Resolution-Adjustable Multimodal Encoder for Earth ObservationCVPR2026RAMENlink
SARMAESARMAE: Masked Autoencoder for SAR Representation LearningCVPR2026SARMAElink

Remote Sensing Vision-Language Foundation Models

AbbreviationTitlePublicationPaperCode & Weights
-Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsArxiv2023Paperlink
-Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual ModelsArxiv2024Paperlink
SkyCLIPSkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingAAAI2024SkyCLIPlink
GeoRSCLIPRS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote SensingIEEE TGRS2024GeoRSCLIPlink
RemoteCLIPRemoteCLIP: A Vision Language Foundation Model for Remote SensingIEEE TGRS2024RemoteCLIPlink
EarthGPTEarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing DomainIEEE TGRS2024EarthGPTlink
GRAFTRemote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentICLR2024GRAFTnull
EarthMarkerEarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote SensingIEEE TGRS2024EarthMarkerlink
RingMoGPTRingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded TasksIEEE TGRS2024RingMoGPTnull
RS-LLaVARS-LLaVA: Large Vision Language Model for Joint Captioning and Question Answering in Remote Sensing ImageryRS2024RS-LLaVAlink
GeoChatGeoChat: Grounded Large Vision-Language Model for Remote SensingCVPR2024GeoChatlink
SkySenseGPTSkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingArxiv2024SkySenseGPTlink
RSCLIPPushing the Limits of Vision-Language Models in Remote Sensing without Human AnnotationsArxiv2024RSCLIPnull
GeoTextTowards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation MatchingECCV2024GeoTextlink
LHRS-BotLHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language ModelECCV2024LHRS-Botlink
GeoGroundGeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual GroundingArxiv2024GeoGroundlink
RSUniVLMRSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of ExpertsArxiv2024RSUniVLMnull
REO-VLMREO-VLM: Transforming VLM to Meet Regression Challenges in Earth ObservationArxiv2024REO-VLMnull
UniRSUniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language ModelsArxiv2024UniRSnull
VHMVHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisAAAI2025VHMlink
-Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsArxiv2025Papernull
DOFA-CLIPDOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth ObservationArxiv2025DOFA-CLIPlink
FalconFalcon: A Remote Sensing Vision-Language Foundation ModelArxiv2025Falconlink
GeoRSMLLMGeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote SensingArxiv2025GeoRSMLLMnull
OmniGeoOmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial IntelligenceArxiv2025OmniGeonull
DGTRS-CLIPDGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for AlignmentArxiv2025DGTRS-CLIPlink
EagleVisionEagleVision: Object-level Attribute Multimodal LLM for Remote SensingArxiv2025EagleVisionlink
SkyEyeGPTSkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language ModelISPRS JPRS2025SkyEyeGPTlink
SegEarth-R1SegEarth-R1: Geospatial Pixel Reasoning via Large Language ModelArxiv2025SegEarth-R1link
EarthGPT-XEarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual PromptingArxiv2025EarthGPT-Xlink
TEOChatTEOChat: A Large Vision-Language Assistant for Temporal Earth Observation DataICLR2025TEOChatlink
LISAtLISAt: Language-Instructed Segmentation Assistant for Satellite ImageryArxiv2025LISAtnull
DynamicVLDynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City UnderstandingArxiv2025DynamicVLnull
RSGPTRSGPT: A Remote Sensing Vision Language Model and BenchmarkISPRS JPRS2025RSGPTlink
LHRS-Bot-NovaLHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language InterpretationISPRS JPRS2025LHRS-Bot-Novalink
EarthDialEarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesCVPR2025EarthDiallink
SkySense-OSkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingCVPR2025SkySense-Olink
XLRS-BenchXLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?CVPR2025XLRS-Benchnull
GeoPixGeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote SensingIEEE GRSM2025GeoPixlink
Co-LLaVACo-LLaVA: Efficient Remote Sensing Visual Question Answering via Model CollaborationRS2025Co-LLaVAnull
RLitaRLita: A Region-Level Image-Text Alignment Method for Remote Sensing Foundation ModelRS2025RLitanull
EarthMindEarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLMArxiv2025EarthMindlink
-Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert ModelingArxiv2025Papernull
GeoPixelGeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingICML2025GeoPixellink
RingMo-AgentRingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal ReasoningArxiv2025RingMo-Agentnull
Geo-R1Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-TuningArxiv2025Geo-R1null
RSThinkerTowards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language ModelsArxiv2025RSThinkernull
FUSAR-KLIPFUSAR-KLIP: Towards Multimodal Foundation Models for Remote SensingArxiv2025FUSAR-KLIPlink
GeoMagGeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image ParsingACMMM2025GeoMagnull
RemoteSAMRemoteSAM: Towards Segment Anything for Earth ObservationACMMM2025RemoteSAMlink
LRS-VQAWhen Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token PruningICCV2025LRS-VQAlink
UrbanLLaVAUrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and UnderstandingICCV2025UrbanLLaVAlink
SARCLIPSARCLIP: a multimodal foundation framework for SAR imagery via contrastive language-image pre-trainingISPRS JPRS2025SARCLIPlink
FUSE-RSVLMFUSE-RSVLM: Feature Fusion Vision-Language Model for Remote SensingArxiv2025FUSE-RSVLMlink
AquilaAquila: A Hierarchically Aligned Vision-Language Model for Enhanced Remote Sensing Image ComprehensionJournal of Remote Sensing 2026Aquilanull
RSCoVLMCo-Training Vision-Language Models for Remote Sensing Multi-Task LearningRS2026RSCoVLMlink
GeoReasonGeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement LearningArxiv2026GeoReasonlink
RemoteReasonerRemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowAAAI2026RemoteReasonernull
SkyMoESkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of ExpertsAAAI2026SkyMoEnull
GeoEyesGeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing ImageryArxiv2026GeoEyeslink
GeoSolverGeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process SupervisionArxiv2026GeoSolvernull
GeoAlignCLIPGeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency LearningArxiv2026GeoAlignCLIPnull
RS-WorldModelRS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense ForecastingArxiv2026RS-WorldModelnull
Decoding-the-DeltaDecoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language ModelsArxiv2026Papernull
RemoteShieldRemoteShield: Enable Robust Multimodal Large Language Models for Earth ObservationArxiv2026RemoteShieldnull
UniChangeUniChange: Unifying Change Detection with Multimodal Large Language ModelCVPR2026UniChangelink
ZoomEarthZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language TasksCVPR2026ZoomEarthnull
UniGeoSegUniGeoSeg: Towards Unified Open-World Segmentation for Geospatial ScenesCVPR2026UniGeoSeglink
SegEarth-R2SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesCVPR2026SegEarth-R2link
FUSAR-GPTFUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR ImageryCVPR2026FUSAR-GPTnull
SATtxtSpectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite ImageryCVPR2026SATtxtlink
TerraScopeTerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationCVPR2026TerraScopenull

Remote Sensing Generative Foundation Models

AbbreviationTitlePublicationPaperCode & Weights
GeoRSSDRS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation ModelArxiv2023Paperlink
-Generate Your Own Scotland: Satellite Image Generation Conditioned on MapsNeurIPSW2023Paperlink
CRS-DiffCRS-Diff: Controllable Remote Sensing Image Generation with Diffusion ModelArxiv2024Paperlink
DiffusionSatDiffusionSat: A Generative Foundation Model for Satellite ImageryICLR2024DiffusionSatlink
HSIGeneHSIGene: A Foundation Model For Hyperspectral Image GenerationArxiv2024Paperlink
Text2EarthText2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation ModelGRSM2025Text2Earthlink
MetaEarthMetaEarth: A Generative Foundation Model for Global-Scale Remote Sensing Image GenerationTPAMI2025MetaEarthlink
EcoMapperEcoMapper: Generative Modeling for Climate-Aware Satellite ImageryICML2025EcoMapperlink
OSMGenOSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap DataNeurIPSW2025OSMGenlink
Any2RSIAny2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched DescriptionAAAI2026Any2RSIlink
GeoDiTGeoDiT: Point-Conditioned Diffusion Transformer for Satellite Image SynthesisArxiv2026GeoDiTnull
MetaEarth3DMetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative ModelingArxiv2026MetaEarth3Dnull

Remote Sensing Vision-Location Foundation Models

AbbreviationTitlePublicationPaperCode & Weights
CSPCSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual RepresentationsICML2023CSPlink
GeoCLIPGeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationNeurIPS2023GeoCLIPlink
SatCLIPSatCLIP: Global, General-Purpose Location Embeddings with Satellite ImageryAAAI2025SatCLIPlink
GAIRGAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit RepresentationsArxiv2025GAIRlink
RANGERANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsCVPR2025RANGEnull
Geo²Geo²: Geometry-Guided Cross-view Geo-Localization and Image SynthesisCVPR2026Geo²null
GeoBridgeGeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-LocalizationCVPR2026GeoBridgelink

Remote Sensing Vision-Audio Foundation Models

AbbreviationTitlePublicationPaperCode & Weights
-Self-supervised audiovisual representation learning for remote sensing dataJAG2022Paperlink
GeoBindGeoBind: Binding Text, Image, and Audio through Satellite ImagesIGARSS2024GeoBindnull
PSMPSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingACM MM 2024PSMlink
Sat2SoundSat2Sound: A Unified Framework for Zero-Shot Soundscape MappingArxiv2025Sat2Soundnull

Remote Sensing Agents

AbbreviationTitlePublicationPaperCode & Weights
GeoLLM-QAEvaluating Tool-Augmented Agents in Remote Sensing PlatformsICLR 2024 ML4RS WorkshopPapernull
GeoLLM-EngineGeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots.CVPRW2024Papernull
RS-AgentRS-Agent: Automating Remote Sensing Tasks through Intelligent AgentArxiv2024Papernull
Change-AgentChange-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and AnalysisTGRS2024Paperlink
GeoLLM-SquadMulti-Agent Geospatial Copilots for Remote Sensing WorkflowsArxiv2025Papernull
-Towards LLM Agents for Earth Observation: The UnivEARTH DatasetArxiv2025Papernull
AirSpatialBotAirSpatialBot: A Spatially Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognition and RetrievalIEEE TGRS2025Paperlink
ThinkGeoThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing TasksArxiv2025Paperlink
PEACEPEACE: Empowering Geologic Map Holistic Understanding with MLLMsCVPR2025Paperlink
Geo-OLMGeo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven WorkflowsCOMPASS'2025Paperlink
REMSAREMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware AgentArxiv2025Paperlink
OpenEarthAgentOpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial AgentsArxiv2026Paperlink
OpenEarth-AgentOpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth ObservationArxiv2026Paperlink
Earth-AgentEarth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsICLR2026Paperlink
GeoMMAgentGeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingCVPR2026Papernull
IMAIAIMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial IntelligenceCVPR2026Papernull

Benchmarks for RSFMs

AbbreviationTitlePublicationPaperLinkDownstream Tasks
-Revisiting pre-trained remote sensing model benchmarks: resizing and normalization mattersArxiv2023PaperlinkClassification
GEO-BenchGEO-Bench: Toward Foundation Models for Earth MonitoringArxiv2023PaperlinkClassification & Segmentation
FoMo-BenchFoMo: Multi-Modal, Multi-Scale and Multi-Task Remote Sensing Foundation Models for Forest MonitoringArxiv2023FoMo-BenchComing soonClassification & Segmentation & Detection for forest monitoring
PhilEOPhilEO Bench: Evaluating Geo-Spatial Foundation ModelsArxiv2024PaperlinkSegmentation & Regression estimation
SkySenseSkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryCVPR2024SkySenseTargeted open-sourceClassification & Segmentation & Detection & Change detection & Multi-Modal Segmentation: Time-insensitive LandCover Mapping & Multi-Modal Segmentation: Time-sensitive Crop Mapping & Multi-Modal Scene Classification
VLEO-BenchGood at captioning, bad at counting: Benchmarking GPT-4V on Earth observation dataArxiv2024VLEO-benchlinkLocation Recognition & Captioning & Scene Classification & Counting & Detection & Change detection
VRSBenchVRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image UnderstandingNeurIPS2024VRSBenchlinkImage Captioning & Object Referring & Visual Question Answering
UrBenchUrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosAAAI2025UrBenchlinkObject Referring & Visual Question Answering & Counting & Scene Classification & Location Recognition & Geolocalization
PANGAEAPANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation ModelsArxiv2024PANGAEAlinkSegmentation & Change detection & Regression
CHOICECHOICE: Evaluating and Understanding Vision-Language Model Choices in Remote SensingNeurIPS2025CHOICElinkPerception & Reasoning
GEO-Bench-VLMGEO-Bench-VLM: Benchmarking Vision-Language Models for Geospatial TasksICCV2025GEO-Bench-VLMlinkScene Understanding & Counting & Object Classification & Event Detection & Spatial Relations
Copernicus-BenchTowards a Unified Copernicus Foundation Model for Earth VisionICCV2025Copernicus-BenchlinkSegmentation & Classification & Change detection & Regression
REOBenchREOBench: Benchmarking Robustness of Earth Observation Foundation ModelsArxiv2025REOBenchlinkRobustness across 6 Earth observation tasks
Plantation BenchPlantation Bench: A Multiscale, Multimodal Remote Sensing Benchmark for Plantation Mapping Under Distribution ShiftICCVW2025Plantation BenchnullPlantation Mapping under Distribution Shift
ChatEarthBenchChatEarthBench: Benchmarking multimodal large language models for Earth observationIEEE GRSM2026ChatEarthBenchnullBenchmarking EO multimodal large language models
GeoReason-BenchGeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement LearningArxiv2026GeoReason-BenchlinkLogical consistency & multi-step reasoning
Earth-BenchEarth-Agent: Unlocking the Full Landscape of Earth Observation with Agents (Earth-Bench is the benchmark introduced in this paper)ICLR2026Earth-AgentlinkTool-augmented EO reasoning & multi-step planning & quantitative spatiotemporal analysis
OmniEarthOmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial TasksArxiv2026OmniEarthlinkPerception & Reasoning & Robustness across geospatial tasks
SpatialSky-BenchIs your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV NavigationCVPR2026SpatialSky-BenchnullUAV spatial intelligence evaluation for Vision-Language Models
GeoMMBenchGeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingCVPR2026GeoMMBenchnullExpert-level multimodal QA across RS disciplines, sensors & tasks
RSVLM-QARSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question AnsweringACM MM2025RSVLM-QAlinkVisual Question Answering & Image Captioning & Counting
Geo3DVQAGeo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial ImageryWACV2026Geo3DVQAlink3D geospatial reasoning & height-aware spatial analysis
OpenEarth-BenchOpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation (OpenEarth-Bench is the benchmark introduced in this paper)Arxiv2026OpenEarth-AgentlinkFull-pipeline EO across 7 application domains
GeoAgentBenchGeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial AnalysisArxiv2026GeoAgentBenchnullDynamic execution evaluation for tool-augmented GIS agents

(Large-scale) Pre-training Datasets

AbbreviationTitlePublicationPaperAttributeLink
fMoWFunctional Map of the WorldCVPR2018fMoWVisionlink
SEN12MSSEN12MS -- A Curated Dataset of Georeferenced Multi-Spectral Sentinel-1/2 Imagery for Deep Learning and Data Fusion-SEN12MSVisionlink
BEN-MMBigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and RetrievalGRSM2021BEN-MMVisionlink
MillionAIDOn Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances, and Million-AIDJSTARS2021MillionAIDVisionlink
SeCoSeasonal Contrast: Unsupervised Pre-Training From Uncurated Remote Sensing DataICCV2021SeCoVisionlink
fMoW-S2SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryNeurIPS2022fMoW-S2Visionlink
TOV-RS-BalancedTOV: The original vision model for optical remote sensing image understanding via self-supervised learningJSTARS2023TOVVisionlink
SSL4EO-S12SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth ObservationGRSM2023SSL4EO-S12Visionlink
SSL4EO-LSSL4EO-L: Datasets and Foundation Models for Landsat ImageryArxiv2023SSL4EO-LVisionlink
SatlasPretrainSatlasPretrain: A Large-Scale Dataset for Remote Sensing Image UnderstandingICCV2023SatlasPretrainVision (Supervised)link
CACoChange-Aware Sampling and Contrastive Learning for Satellite ImagesCVPR2023CACoVisionComing soon
SAMRSSAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything ModelNeurIPS2023SAMRSVisionlink
RSVGRSVG: Exploring Data and Models for Visual Grounding on Remote Sensing DataTGRS2023RSVGVision-Languagelink
RS5MRS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation ModelArxiv2023RS5MVision-Languagelink
GEO-BenchGEO-Bench: Toward Foundation Models for Earth MonitoringArxiv2023GEO-BenchVision (Evaluation)link
RSICap & RSIEvalRSGPT: A Remote Sensing Vision Language Model and BenchmarkArxiv2023RSGPTVision-LanguageComing soon
ClayClay Foundation Model-nullVisionlink
SATINSATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language ModelsICCVW2023SATINVision-Languagelink
SkyScriptSkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingAAAI2024SkyScriptVision-Languagelink
ChatEarthNetChatEarthNet: a global-scale image-text dataset empowering vision-language geo-foundation modelsESSD2025ChatEarthNetVision-Languagelink
LuoJiaHOGLuoJiaHOG: A hierarchy oriented geo-aware image caption dataset for remote sensing image-text retrievalISPRS JPRS2025LuoJiaHOGVision-Languagenull
MMEarthMMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation LearningArxiv2024MMEarthVisionlink
SeeFarSeeFar: Satellite Agnostic Multi-Resolution Dataset for Geospatial Foundation ModelsArxiv2024SeeFarVisionlink
FIT-RSSkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingArxiv2024PaperVision-Languagelink
RS-GPT4VRS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image UnderstandingArxiv2024PaperVision-Languagelink
RS-4MHarnessing Massive Satellite Imagery with Efficient Masked Image ModelingICCV2025RS-4MVisionlink
Major TOMMajor TOM: Expandable Datasets for Earth ObservationArxiv2024Major TOMVisionlink
VRSBenchVRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image UnderstandingNeurIPS2024VRSBenchVision-Languagelink
MMM-RSMMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image GenerationArxiv2024MMM-RSVision-Languagelink
DDFAVDDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation BenchmarkRS2025DDFAVVision-Languagelink
M3LEOA Multi-Modal, Multi-Label Earth Observation Dataset Integrating Interferometric SAR and Multispectral DataNeurIPS2024M3LEOVisionlink
Copernicus-PretrainTowards a Unified Copernicus Foundation Model for Earth VisionICCV2025Copernicus-PretrainVisionlink
DGTRSDDGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for AlignmentArxiv2025PaperVision-Languagelink
EarthDial-InstructEarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesCVPR2025PaperVision-Languagelink
GeoPixelDGeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingICML2025PaperVision-Languagelink
GeoPixInstructGeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote SensingIEEE GRSM2025PaperVision-Languagelink
GeoLangBind-2MRethinking Remote Sensing CLIP: Leveraging Multimodal Large Language Models for High-Quality Vision-Language DatasetICONIP2024PaperVision-Languagelink
Falcon_SFTFalcon: A Remote Sensing Vision-Language Foundation ModelArxiv2025PaperVision-Languagelink
UnivEARTHTowards LLM Agents for Earth Observation: The UnivEARTH DatasetArxiv2025PaperVision-Language & Agentsnull
RemoteSAM-270KRemoteSAM: Towards Segment Anything for Earth ObservationACMMM2025PaperVision-Languagelink
OpenEarthAgent DatasetOpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial AgentsArxiv2026PaperVision-Language & Agentslink
UHR-CoZGeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing ImageryArxiv2026PaperVision-Languagelink
SOMA-1MSOMA-1M: A Large-Scale SAR-Optical Multi-resolution Alignment Dataset for Multi-Task Remote SensingArxiv2026SOMA-1MVision (SAR-Optical)link

Embeddings data

AbbreviationTitlePublicationPaperCodeDataset / Product
CLAY EmbeddingsClay Model v0 EmbeddingsSource Cooperative2024nulllinklink
Major TOM EmbeddingsGlobal and Dense Embeddings of Earth: Major TOM Floating in the Latent SpaceArxiv2024Paperlinklink
Earth Genome EmbeddingsEmbeddings for allMedium2025Papernulllink
TESSERATESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisCVPR2026Paperlinklink
AlphaEarthAlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label dataArxiv2025Papernulllink
ESDDemocratizing planetary-scale analysis: An ultra-lightweight Earth embedding database for accurate and flexible global land monitoringArxiv2026Paperlinklink
Copernicus-EmbedTowards a Unified Copernicus Foundation Model for Earth VisionICCV2025 OralPaperlinklink

Relevant Projects

TitleLinkBrief Introduction
RSFMs (Remote Sensing Foundation Models) PlaygroundlinkAn open-source playground to streamline the evaluation and fine-tuning of RSFMs on various datasets.
PANGAEAlinkA Global and Inclusive Benchmark for Geospatial Foundation Models.
GeoFMlinkEvaluation of Foundation Models for Earth Observation.
rs-embedlinkOne line code to get Any Remote Sensing Foundation Model (RSFM) embeddings for Any Place and Any Time.
TerraTorchlinkA PyTorch toolkit for fine-tuning Geospatial Foundation Models, supporting Prithvi, TerraMind, SatMAE, ScaleMAE, DOFA, Clay, and more.
TorchGeolinkA PyTorch domain library for geospatial data, providing datasets, samplers, transforms, and pre-trained models.
Awesome-Geospatial-EmbeddingslinkA curated list of papers on how to represent Earth data in embedding space — spatial, temporal, or semantic.

Survey/Commentary Papers

TitlePublicationPaperAttribute
The Potential of Visual ChatGPT For Remote SensingArxiv2023PaperVision-Language
Self-Supervised Remote Sensing Feature Learning: Learning Paradigms, Challenges, and Future WorksTGRS2023PaperVision & Vision-Language
Revisiting pre-trained remote sensing model benchmarks: resizing and normalization mattersArxiv2023PaperVision
An Agenda for Multimodal Foundation Models for Earth ObservationIGARSS2023PaperVision
遥感大模型:进展与前瞻武汉大学学报 (信息科学版) 2023PaperVision & Vision-Language
地理人工智能样本:模型、质量与服务武汉大学学报 (信息科学版) 2023Paper-
Brain-Inspired Remote Sensing Foundation Models and Open Problems: A Comprehensive SurveyJSTARS2023PaperVision & Vision-Language
遥感基础模型发展综述与未来设想遥感学报2023Paper-
On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning ApplicationsArxiv2023PaperVision-Language
Transfer learning in environmental remote sensingRSE2024PaperTransfer learning
Vision-Language Models in Remote Sensing: Current Progress and Future TrendsIEEE GRSM2024PaperVision-Language
On the Foundations of Earth and Climate Foundation ModelsArxiv2024PaperVision & Vision-Language
多模态遥感基础大模型:研究现状与未来展望测绘学报2024PaperVision & Vision-Language & Generative & Vision-Location
Towards Vision-Language Geo-Foundation Model: A SurveyArxiv2024PaperVision-Language
Vision Foundation Models in Remote Sensing: A SurveyArxiv2024PaperVision
Foundation model for generalist remote sensing intelligence: Potentials and prospectsScience Bulletin2024Paper-
Advancements in Visual Language Models for Remote Sensing: Datasets, Capabilities, and Enhancement TechniquesArxiv2024PaperVision-Language
When Geoscience Meets Foundation Models: Toward a general geoscience artificial intelligence systemIEEE GRSM2024PaperVision & Vision-Language
Towards the next generation of Geospatial Artificial IntelligenceJAG2025Paper-
When Remote Sensing Meets Foundation Model: A Survey and BeyondRS2025PaperVision & Vision-Language & Generative & Agents
Vision Foundation Models in Remote Sensing: A surveyIEEE GRSM2025PaperVision
Remote Sensing Tuning: A SurveyCVM2025PaperVision & Vision-Language
Unleashing the potential of remote sensing foundation models via bridging data and computility islandsThe Innovation2025Paper-
A Survey on Remote Sensing Foundation Models: From Vision to MultimodalityArxiv2025Paper-
Foundation Models for Remote Sensing and Earth Observation: A surveyIEEE GRSM2025PaperVision & Vision-Language
Vision-Language Modeling Meets Remote Sensing: Models, datasets, and perspectivesIEEE GRSM2025PaperVision-Language
MIMRS: A Survey on Masked Image Modeling in Remote SensingIGARSS2025PaperVision
A Review of Challenges and Applications in Remote Sensing Foundation ModelsIGARSS2025PaperVision & Vision-Language
On the Status of Foundation Models for SAR ImageryArxiv2025PaperVision (SAR)
Advances on Multimodal Remote Sensing Foundation Models for Earth Observation Downstream Tasks: A SurveyRS2025PaperVision & Vision-Language
Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging SystemsWACVW2026PaperAgents
Onboard Deployment of Remote Sensing Foundation Models: A Comprehensive Review of Architecture, Optimization, and HardwareRS2026PaperVision & Vision-Language
On the foundations of Earth foundation modelsCommunications Earth & Environment 2026PaperVision & Vision-Language
A Genealogy of Foundation Models in Remote SensingACM TSAS2026PaperVision & Vision-Language
Foundation Models in Remote Sensing: Evolving from Unimodality to MultimodalityIEEE GRSM2026PaperVision & Vision-Language

Citation

If you find this repository useful, please consider giving a star :star: and citation:

@inproceedings{guo2024skysense,
  title={Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery},
  author={Guo, Xin and Lao, Jiangwei and Dang, Bo and Zhang, Yingying and Yu, Lei and Ru, Lixiang and Zhong, Liheng and Huang, Ziyuan and Wu, Kang and Hu, Dingxiang and others},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={27672--27683},
  year={2024}
}

@article{li2025unleashing,
  title={Unleashing the potential of remote sensing foundation models via bridging data and computility islands},
  author={Li, Yansheng and Tan, Jieyi and Dang, Bo and Ye, Mang and Bartalev, Sergey A and Shinkarenko, Stanislav and Wang, Linlin and Zhang, Yingying and Ru, Lixiang and Guo, Xin and others},
  journal={The Innovation},
  year={2025},
  publisher={Elsevier}
}

@article{wu2025semantic,
  author = {Wu, Kang and Zhang, Yingying and Ru, Lixiang and Dang, Bo and Lao, Jiangwei and Yu, Lei and Luo, Junwei and Zhu, Zifan and Sun, Yue and Zhang, Jiahao and Zhu, Qi and Wang, Jian and Yang, Ming and Chen, Jingdong and Zhang, Yongjun and Li, Yansheng},
  title= {A semantic‑enhanced multi‑modal remote sensing foundation model for Earth observation},
  journal= {Nature Machine Intelligence},
  year= {2025},
  doi= {10.1038/s42256-025-01078-8},
  url= {https://doi.org/10.1038/s42256-025-01078-8}
}

@inproceedings{zhu2025skysense,
  title={Skysense-o: Towards open-world remote sensing interpretation with vision-centric visual-language modeling},
  author={Zhu, Qi and Lao, Jiangwei and Ji, Deyi and Luo, Junwei and Wu, Kang and Zhang, Yingying and Ru, Lixiang and Wang, Jian and Chen, Jingdong and Yang, Ming and others},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={14733--14744},
  year={2025}
}

@article{luo2024skysensegpt,
  title={Skysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understanding},
  author={Luo, Junwei and Pang, Zhen and Zhang, Yongjun and Wang, Tingzhu and Wang, Linlin and Dang, Bo and Lao, Jiangwei and Wang, Jian and Chen, Jingdong and Tan, Yihua and others},
  journal={arXiv preprint arXiv:2406.10100},
  year={2024}
}