๐Ÿ”— Links and References

October 7, 2025 ยท View on GitHub

Links to research papers and resources corresponding to implemented features in this repository. Please feel free to fill in any missing references!

OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics introduces

  • The technical report on OpenUnlearning, its design, features and other details.
  • A meta-evaluation framework to benchmark unlearning evaluations on a set of 450+ open sourced models.
  • Results benchmarking 8 diverse unlearning methods in one place using 10 evaluation metrics on TOFU.

๐Ÿ“Œ Table of Contents


๐Ÿ“— Implemented Methods

MethodResource
GradAscent, GradDiffNaive baselines found in many papers including MUSE, TOFU etc.
NPOPaper๐Ÿ“„, Code ๐Ÿ™
SimNPOPaper๐Ÿ“„, Code ๐Ÿ™
IdkDPOTOFU (๐Ÿ“„)
RMUWMDP paper (๐Ÿ™, ๐ŸŒ), later used in G-effect (๐Ÿ™)
UNDIALPaper๐Ÿ“„, Code ๐Ÿ™
AltPOPaper๐Ÿ“„, Code ๐Ÿ™
SatImpPaper๐Ÿ“„, Code ๐Ÿ™
WGA (G-effect)Paper๐Ÿ“„, Code ๐Ÿ™
CE-U (Cross-Entropy unlearning)Paper๐Ÿ“„
PDUPaper ๐Ÿ“„

๐Ÿ“˜ Benchmarks

BenchmarkResource
TOFUPaper๐Ÿ“„
MUSEPaper๐Ÿ“„
WMDPPaper๐Ÿ“„

๐Ÿ“™ Evaluation Metrics

MetricResource
Verbatim Probability / ROUGE, simple QA-ROUGENaive metrics found in many papers including MUSE, TOFU etc.
Membership Inference Attacks (LOSS, ZLib, Reference, GradNorm, MinK, MinK++)MIMIR (๐Ÿ™), MUSE (๐Ÿ“„)
PrivLeakMUSE (๐Ÿ“„)
Forget Quality, Truth Ratio, Model UtilityTOFU (๐Ÿ“„)
Extraction Strength (ES)Carlini et al., 2021 (๐Ÿ“„), used for unlearning in Wang et al., 2025 (๐Ÿ“„)
Exact Memorization (EM)Tirumala et al., 2022 (๐Ÿ“„), used for unlearning in Wang et al., 2025 (๐Ÿ“„)
lm-evaluation-harnessRepository: ๐Ÿ’ป

๐Ÿ“š Surveys

๐Ÿ™ Other GitHub Repositories