README.md
September 8, 2026 · View on GitHub
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
Tadej Tomanič1, 2 · Alice Baudhuin1 · Jan Sotošek1 · Jure Brence1, 3 · Panče Panov1, 3 · Nikola Simidjievski1, 4 · Dragi Kocev1, 3
1Bias Variance Labs, d.o.o. 2University of Ljubljana, Faculty of Mathematics and Physics
3Department of Knowledge Technologies, Jožef Stefan Institute 4Télécom Paris, Institut Polytechnique de Paris
Summary
Code and experiments for the paper, A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation, by Tadej Tomanič, Alice Baudhuin, Jan Sotošek, Jure Brence, Panče Panov, Nikola Simidjievski, and Dragi Kocev (currently under review).
This study presents a standardized benchmark for change detection in Earth Observation, developed as part of the OSCARS-funded FAIR-EO project. The framework is integrated into the AiTLAS toolbox.
Despite rapid advancements in machine learning for remote sensing, accurately evaluating and comparing change detection models remains a significant challenge. Variations in dataset preprocessing, data splits, and evaluation metrics often lead to inconsistent results, making it difficult to determine whether a new method genuinely outperforms existing ones. To address this, we introduce a comprehensive, transparent, and trustworthy benchmarking framework designed to eliminate these ambiguities.
Aligned with the core principles of FAIR (Findable, Accessible, Interoperable, and Reusable) and open science, our framework provides a unified pipeline for the Earth Observation community. It establishes rigorous testing protocols, standardized evaluation metrics, and reproducible baselines.
Datasets
We do not host the datasets, but we provide links to original data sources, and our custom data splits and data loaders used in this study.
| Dataset | Paper | Year | Data source | Data splits | Dataloader |
|---|---|---|---|---|---|
| BANDON | Pang et al. | 2023 | GitHub | Train • Validation • Test | Python file |
| CLCD | Liu et al. | 2022 | GitHub | Train • Validation • Test | Python file |
| DSIFN | Zhang et al. | 2020 | GitHub | Train • Validation • Test | Python file |
| EGY-BCD | Holail et al. | 2023 | GitHub | Train • Validation • Test | Python file |
| LEVIR-CD+ | Chen and Shi | 2020 | GitHub | Train • Validation • Test | Python file |
| MSBC | Li et al. | 2022 | GitHub | Train • Validation • Test | Python file |
| MSOSCD | Li et al. | 2022 | GitHub | Train • Validation • Test | Python file |
| OMBRIA | Drakonakis et al. | 2022 | GitHub | Train • Validation • Test | Python file |
| Season-Varying CDD | Lebedev et al. | 2018 | Google Drive | Train • Validation • Test | Python file |
| SYSU-CD | Shi et al. | 2022 | GitHub | Train • Validation • Test | Python file |
Model checkpoints & TensorBoard logs
Model checkpoints and TensorBoard logs are available on HuggingFace.
| Model | Paper | Year | # Params (M) | GFLOPs | Model checkpoints | TensorBoard logs |
|---|---|---|---|---|---|---|
| BIT | Chen et al. | 2021 | 3.50 | 10.88 | Scratch • Pre-trained | Scratch • Pre-trained |
| CGNet | Han et al. | 2023 | 33.68 | 87.55 | Scratch • Pre-trained | Scratch • Pre-trained |
| ChangeFormerV6 | Bandara and Patel | 2022 | 41.03 | 138.77 | Scratch • Pre-trained | Scratch • Pre-trained |
| ChangeViT | Zhu et al. | 2026 | 20.66 | 26.36 | Scratch • Pre-trained | Scratch • Pre-trained |
| CSSM | Ghazaei et a. | 2025 | 2.59 | 2.78 | Scratch • Pre-trained | Scratch • Pre-trained |
| HRNet SiamConc | Sun et al. | 2019 | 21.48 | 9.55 | Scratch • Pre-trained | Scratch • Pre-trained |
| SiamCRNN | Chen et al. | 2020 | 28.51 | 65.49 | Scratch • Pre-trained | Scratch • Pre-trained |
| STANet | Chen and Shi | 2020 | 12.21 | 19.15 | Scratch • Pre-trained | Scratch • Pre-trained |
| TinyCD | Codegoni et al. | 2023 | 0.29 | 1.46 | Scratch • Pre-trained | Scratch • Pre-trained |
| U-Net SiamConc | Caye Daudt et al. | 2018 | 40.35 | 19.37 | Scratch • Pre-trained | Scratch • Pre-trained |
Performance
Mean Intersection over Union (mIoU) for models trained from scratch. The best performance is indicated in bold, and the second-best is underlined.
| Dataset/Model | BIT | CGNet | ChangeFormerV6 | ChangeViT | CSSM | HRNet SiamConc | SiamCRNN | STANet | TinyCD | U-Net SiamConc | Average |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BANDON | 0.5015 | 0.4785 | 0.5661 | 0.5545 | 0.4787 | 0.5231 | 0.5511 | 0.4921 | 0.5306 | 0.4840 | 0.5160 |
| CLCD | 0.6920 | 0.6593 | 0.6569 | 0.6637 | 0.4596 | 0.6375 | 0.6950 | 0.6120 | 0.6612 | 0.7048 | 0.6442 |
| DSIFN | 0.7711 | 0.8109 | 0.7453 | 0.7919 | 0.6829 | 0.6409 | 0.7863 | 0.7224 | 0.7133 | 0.7593 | 0.7424 |
| EGY-BCD | 0.7783 | 0.8214 | 0.7594 | 0.7770 | 0.7005 | 0.7619 | 0.7929 | 0.7598 | 0.7619 | 0.7762 | 0.7689 |
| LEVIR-CD+ | 0.6378 | 0.6412 | 0.6561 | 0.7181 | 0.5785 | 0.6766 | 0.6907 | 0.5449 | 0.7201 | 0.7201 | 0.6584 |
| MSBC | 0.8613 | 0.8846 | 0.8730 | 0.8644 | 0.6554 | 0.8172 | 0.8729 | 0.7875 | 0.7809 | 0.8699 | 0.8267 |
| MSOSCD | 0.6603 | 0.6748 | 0.6612 | 0.6826 | 0.5330 | 0.6140 | 0.7248 | 0.6094 | 0.6003 | 0.6817 | 0.6442 |
| OMBRIA | 0.6438 | 0.6343 | 0.6934 | 0.6632 | 0.6396 | 0.6123 | 0.6302 | 0.5798 | 0.6138 | 0.6302 | 0.6341 |
| Season-varying CDD | 0.8896 | 0.6578 | 0.8110 | 0.9141 | 0.7507 | 0.8584 | 0.9133 | 0.8296 | 0.8556 | 0.8887 | 0.8369 |
| SYSU-CD | 0.7066 | 0.7646 | 0.7212 | 0.7512 | 0.6909 | 0.6654 | 0.7350 | 0.6709 | 0.6650 | 0.7090 | 0.7080 |
| Average | 0.7142 | 0.7027 | 0.7144 | 0.6170 | 0.6807 | 0.7392 | 0.6008 | 0.6903 | 0.7224 |
Mean Intersection over Union (mIoU) for models with pre-trained backbones.
| Dataset/Model | BIT | CGNet | ChangeViT | HRNet SiamConc | SiamCRNN | STANet | TinyCD | U-Net SiamConc | Average |
|---|---|---|---|---|---|---|---|---|---|
| BANDON | 0.4780 | 0.4623 | 0.5844 | 0.5672 | 0.5957 | 0.4979 | 0.5568 | 0.6114 | 0.5442 |
| CLCD | 0.7104 | 0.5061 | 0.7101 | 0.7182 | 0.7088 | 0.5198 | 0.6821 | 0.7521 | 0.6634 |
| DSIFN | 0.7551 | 0.8396 | 0.8282 | 0.8044 | 0.8459 | 0.8075 | 0.7819 | 0.8569 | 0.8149 |
| EGY-BCD | 0.8065 | 0.7409 | 0.8066 | 0.8433 | 0.8445 | 0.8074 | 0.7619 | 0.8338 | 0.8056 |
| LEVIR-CD+ | 0.6690 | 0.5592 | 0.7317 | 0.7249 | 0.7250 | 0.7031 | 0.7254 | 0.7478 | 0.6983 |
| MSBC | 0.8849 | 0.7223 | 0.8738 | 0.8892 | 0.8950 | 0.8592 | 0.7762 | 0.8943 | 0.8494 |
| MSOSCD | 0.7924 | 0.5820 | 0.7735 | 0.7750 | 0.8076 | 0.7280 | 0.5749 | 0.8169 | 0.7313 |
| OMBRIA | 0.6684 | 0.7051 | 0.6992 | 0.6952 | 0.6903 | 0.6282 | 0.6593 | 0.7114 | 0.6821 |
| Season-varying CDD | 0.8946 | 0.7924 | 0.9243 | 0.9224 | 0.9334 | 0.8964 | 0.8501 | 0.9407 | 0.8943 |
| SYSU-CD | 0.7227 | 0.7279 | 0.7774 | 0.7166 | 0.7508 | 0.6642 | 0.7298 | 0.7656 | 0.7319 |
| Average | 0.7382 | 0.6638 | 0.7709 | 0.7657 | 0.7112 | 0.7098 | 0.7931 |
Citation
If you use this benchmark or code in your research, please cite our paper:
@misc{tomanič2026comprehensivetrustworthybenchmarkai,
title={A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation},
author={Tadej Tomanič and Alice Baudhuin and Jan Sotošek and Jure Brence and Panče Panov and Nikola Simidjievski and Dragi Kocev},
year={2026},
eprint={2608.28247},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={[https://arxiv.org/abs/2608.28247](https://arxiv.org/abs/2608.28247)},
}