README.md

September 8, 2026 · View on GitHub

A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation

Tadej Tomanič1, 2 · Alice Baudhuin1 · Jan Sotošek1 · Jure Brence1, 3 · Panče Panov1, 3 · Nikola Simidjievski1, 4 · Dragi Kocev1, 3


1Bias Variance Labs, d.o.o.  2University of Ljubljana, Faculty of Mathematics and Physics
3Department of Knowledge Technologies, Jožef Stefan Institute  4Télécom Paris, Institut Polytechnique de Paris



Summary

Code and experiments for the paper, A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation, by Tadej Tomanič, Alice Baudhuin, Jan Sotošek, Jure Brence, Panče Panov, Nikola Simidjievski, and Dragi Kocev (currently under review).

This study presents a standardized benchmark for change detection in Earth Observation, developed as part of the OSCARS-funded FAIR-EO project. The framework is integrated into the AiTLAS toolbox.

Despite rapid advancements in machine learning for remote sensing, accurately evaluating and comparing change detection models remains a significant challenge. Variations in dataset preprocessing, data splits, and evaluation metrics often lead to inconsistent results, making it difficult to determine whether a new method genuinely outperforms existing ones. To address this, we introduce a comprehensive, transparent, and trustworthy benchmarking framework designed to eliminate these ambiguities.

Aligned with the core principles of FAIR (Findable, Accessible, Interoperable, and Reusable) and open science, our framework provides a unified pipeline for the Earth Observation community. It establishes rigorous testing protocols, standardized evaluation metrics, and reproducible baselines.

Datasets

We do not host the datasets, but we provide links to original data sources, and our custom data splits and data loaders used in this study.

DatasetPaperYearData sourceData splitsDataloader
BANDONPang et al.2023GitHubTrain • Validation • TestPython file
CLCDLiu et al.2022GitHubTrain • Validation • TestPython file
DSIFNZhang et al.2020GitHubTrain • Validation • TestPython file
EGY-BCDHolail et al.2023GitHubTrain • Validation • TestPython file
LEVIR-CD+Chen and Shi2020GitHubTrain • Validation • TestPython file
MSBCLi et al.2022GitHubTrain • Validation • TestPython file
MSOSCDLi et al.2022GitHubTrain • Validation • TestPython file
OMBRIADrakonakis et al.2022GitHubTrain • Validation • TestPython file
Season-Varying CDDLebedev et al.2018Google DriveTrain • Validation • TestPython file
SYSU-CDShi et al.2022GitHubTrain • Validation • TestPython file

Model checkpoints & TensorBoard logs

Model checkpoints and TensorBoard logs are available on HuggingFace.

ModelPaperYear# Params (M)GFLOPsModel checkpointsTensorBoard logs
BITChen et al.20213.5010.88Scratch • Pre-trainedScratch • Pre-trained
CGNetHan et al.202333.6887.55Scratch • Pre-trainedScratch • Pre-trained
ChangeFormerV6Bandara and Patel202241.03138.77Scratch • Pre-trainedScratch • Pre-trained
ChangeViTZhu et al.202620.6626.36Scratch • Pre-trainedScratch • Pre-trained
CSSMGhazaei et a.20252.592.78Scratch • Pre-trainedScratch • Pre-trained
HRNet SiamConcSun et al.201921.489.55Scratch • Pre-trainedScratch • Pre-trained
SiamCRNNChen et al.202028.5165.49Scratch • Pre-trainedScratch • Pre-trained
STANetChen and Shi202012.2119.15Scratch • Pre-trainedScratch • Pre-trained
TinyCDCodegoni et al.20230.291.46Scratch • Pre-trainedScratch • Pre-trained
U-Net SiamConcCaye Daudt et al.201840.3519.37Scratch • Pre-trainedScratch • Pre-trained

Performance

Mean Intersection over Union (mIoU) for models trained from scratch. The best performance is indicated in bold, and the second-best is underlined.

Dataset/ModelBITCGNetChangeFormerV6ChangeViTCSSMHRNet SiamConcSiamCRNNSTANetTinyCDU-Net SiamConcAverage
BANDON0.50150.47850.56610.55450.47870.52310.55110.49210.53060.48400.5160
CLCD0.69200.65930.65690.66370.45960.63750.69500.61200.66120.70480.6442
DSIFN0.77110.81090.74530.79190.68290.64090.78630.72240.71330.75930.7424
EGY-BCD0.77830.82140.75940.77700.70050.76190.79290.75980.76190.77620.7689
LEVIR-CD+0.63780.64120.65610.71810.57850.67660.69070.54490.72010.72010.6584
MSBC0.86130.88460.87300.86440.65540.81720.87290.78750.78090.86990.8267
MSOSCD0.66030.67480.66120.68260.53300.61400.72480.60940.60030.68170.6442
OMBRIA0.64380.63430.69340.66320.63960.61230.63020.57980.61380.63020.6341
Season-varying CDD0.88960.65780.81100.91410.75070.85840.91330.82960.85560.88870.8369
SYSU-CD0.70660.76460.72120.75120.69090.66540.73500.67090.66500.70900.7080
Average0.71420.70270.71440.73810.61700.68070.73920.60080.69030.7224

Mean Intersection over Union (mIoU) for models with pre-trained backbones.

Dataset/ModelBITCGNetChangeViTHRNet SiamConcSiamCRNNSTANetTinyCDU-Net SiamConcAverage
BANDON0.47800.46230.58440.56720.59570.49790.55680.61140.5442
CLCD0.71040.50610.71010.71820.70880.51980.68210.75210.6634
DSIFN0.75510.83960.82820.80440.84590.80750.78190.85690.8149
EGY-BCD0.80650.74090.80660.84330.84450.80740.76190.83380.8056
LEVIR-CD+0.66900.55920.73170.72490.72500.70310.72540.74780.6983
MSBC0.88490.72230.87380.88920.89500.85920.77620.89430.8494
MSOSCD0.79240.58200.77350.77500.80760.72800.57490.81690.7313
OMBRIA0.66840.70510.69920.69520.69030.62820.65930.71140.6821
Season-varying CDD0.89460.79240.92430.92240.93340.89640.85010.94070.8943
SYSU-CD0.72270.72790.77740.71660.75080.66420.72980.76560.7319
Average0.73820.66380.77090.76570.77970.71120.70980.7931

Citation

If you use this benchmark or code in your research, please cite our paper:

@misc{tomanič2026comprehensivetrustworthybenchmarkai,
      title={A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation}, 
      author={Tadej Tomanič and Alice Baudhuin and Jan Sotošek and Jure Brence and Panče Panov and Nikola Simidjievski and Dragi Kocev},
      year={2026},
      eprint={2608.28247},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={[https://arxiv.org/abs/2608.28247](https://arxiv.org/abs/2608.28247)}, 
}