Benchmarking PDFM in Population Modelling
May 1, 2026 ยท View on GitHub
This repository contains the figures and supporting materials for a study benchmarking Population Dynamics Foundation Model (PDFM) embeddings for small-area population modelling.
The study evaluates whether high-dimensional PDFM embeddings can improve population prediction compared with commonly used geospatial covariates. Analyses were conducted across Brazil, Nigeria, and the United States, using administrative-unit population data and multiple modelling approaches.
Repository contents
This repository mainly includes publication-ready figures and related outputs from the PDFM population modelling analysis.
Typical contents include:
- main figures generated for the study
Underlying data/: input data used to generate selected figures, where availableSupplementary data/: processed tables and figure-supporting databoundary/: spatial boundary files used for mapping- R scripts: code used to process results, generate tables, and produce figures
Study overview
The analysis compares PDFM embeddings with conventional geospatial covariates for population modelling. The main objectives are to:
- Benchmark PDFM against conventional geospatial covariates across countries and models.
- Evaluate geographic transferability across held-out regions.
- Assess whether PDFM and geospatial covariates provide complementary information.
- Visualise regional variation in model performance and prediction error.
Models used in the analysis include Random Forest, XGBoost, and Elastic Net. Model performance was evaluated using metrics such as coefficient of determination (R2) and Kullback-Leibler divergence on population shares.
Reproducing figures
To reproduce figures locally, open the relevant R script and update the root directory path