README.md

August 20, 2026 · View on GitHub

Project Logo

PyPI version Docs Python versions License

SafetyCage is a Python package for detecting misclassified predictions from machine learning models in classification tasks. It provides a unified interface for multiple statistical detection methods, enabling users to quantify prediction reliability and flag potentially incorrect outputs across different models and datasets easily.

Available on PyPI: https://pypi.org/project/safetycage/.

Full documentation is available at https://safetycage.readthedocs.io/.

Background

The idea behind safetycage is that we can find statistics on each predicted sample and compare that statistic to some statistic threshold to predict whether the sample prediction was incorrectly classified.

Description

Machine learning models can produce incorrect predictions with high confidence. SafetyCage addresses this by providing post-hoc misclassification detection methods that operate on model outputs or internal representations.

The package includes several methods:

  • MSP (Maximum Softmax Probability)
  • DOCTOR (Error probability estimation)
  • Mahalanobis (Distance-based statistical testing)
  • SPARDACUS (Projection + density estimation approach)

Each method outputs a statistic or p-value that reflects how likely a prediction is to be incorrect.

Alternatively, you can implement your own method by initializing a base class from the safetycage abstract base class, which defines how methods should be implemented.

Requirements

safetycage requires Python 3.13 or later.

Core dependencies (installed automatically): joblib, matplotlib, numpy.

Installation

pip install safetycage
# or: uv add safetycage

Some methods need extra dependencies, installed via pip install safetycage[extra] (or uv add "safetycage[extra]"):

ExtraUsesAdds
redREDtorch, gpytorch
spardacusSPARDACUSstatsmodels, scipy, scikit-learn, tqdm
mahalanobisMahalanobisstatsmodels, scipy
torchTorchModelModuletorch

Tutorials & Examples

To learn how to use safetycage, check out the examples/ directory in this repository. It contains complete integrations with runnable notebooks, including scripts to train models to test the safetycage methods on.

Changelog

See the CHANGELOG.MD for details on versioning.

Support

If you encounter issues or have questions:

Contributing

If you would like to contribute, please reach out to our safetycage team, listed below!

Authors

Acknowledgment

The MSP method was introduced by Hendrycks and Gimpel in A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks.

The DOCTOR method was introduced by Granese et al. in DOCTOR: A Simple Method for Detecting Misclassification Errors.

The Mahalanobis method is described in Johnsen et al..

The SPARDACUS method is described in Johnsen et al..

A proper citation for these methods is provided in the docstring of the code using these methods.

A special thank you goes to previous co-authors of the methods we have built, Filippo Remonato, Shawn Benedict, and Albert Ndur-Osei.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Project status

Active and under development!

Citation

If you use safetycage, please cite us!