Benchmarking PDFM in Population Modelling

May 1, 2026 ยท View on GitHub

This repository contains the figures and supporting materials for a study benchmarking Population Dynamics Foundation Model (PDFM) embeddings for small-area population modelling.

The study evaluates whether high-dimensional PDFM embeddings can improve population prediction compared with commonly used geospatial covariates. Analyses were conducted across Brazil, Nigeria, and the United States, using administrative-unit population data and multiple modelling approaches.

Repository contents

This repository mainly includes publication-ready figures and related outputs from the PDFM population modelling analysis.

Typical contents include:

  • main figures generated for the study
  • Underlying data/: input data used to generate selected figures, where available
  • Supplementary data/: processed tables and figure-supporting data
  • boundary/: spatial boundary files used for mapping
  • R scripts: code used to process results, generate tables, and produce figures

Study overview

The analysis compares PDFM embeddings with conventional geospatial covariates for population modelling. The main objectives are to:

  1. Benchmark PDFM against conventional geospatial covariates across countries and models.
  2. Evaluate geographic transferability across held-out regions.
  3. Assess whether PDFM and geospatial covariates provide complementary information.
  4. Visualise regional variation in model performance and prediction error.

Models used in the analysis include Random Forest, XGBoost, and Elastic Net. Model performance was evaluated using metrics such as coefficient of determination (R2) and Kullback-Leibler divergence on population shares.

Figure1

Reproducing figures

To reproduce figures locally, open the relevant R script and update the root directory path