Published numbers

August 13, 2026 · View on GitHub

A number a reader cannot reproduce is not evidence. This table is the whole reader-facing surface's quantitative content, each figure against the exact command that produces it.

Scope, stated so a green result cannot be over-read. The audit enforces registration on README.md, MCP_LISTINGS.md and index.html: every numeric token a reader sees on those three must be a row below or a declared non-claim, and python claims_audit.py --numbers fails otherwise. CHANGELOG.md and docs/ are not token-enforced — they are covered only by the weaker check that every artifact path they name exists. Their numbers are not audited here, and reading this page as "every number in the project is backed" would be exactly the over-read it exists to prevent.

The ratio

  • 350 numeric tokens are published across the 6 enforced files: README.md, docs/DEEP_DIVE.md, MCP_LISTINGS.md, index.html, compare.html, claude-code.html.
  • 190 of those are quantitative claims, in 103 registry rows below.
  • 78 rows (78/103) are reproducible by a command committed to this repository (REPRODUCIBLE needs nothing but this checkout; REPRODUCIBLE-WITH-DEPS needs a service or dataset we cannot redistribute, named in the command column).
  • The remaining 25 are PENDING-HARNESS, EXTERNAL or WITHDRAWN.
  • The other 160 tokens are declared non-claims — citation years, article numbers, ordinals, ports, example literals — each with a reason and an exact expected count, so adding one silently is not possible either.

Counts by status:

  • REPRODUCIBLE — 26
  • REPRODUCIBLE-WITH-DEPS — 52
  • PENDING-HARNESS — 2
  • EXTERNAL — 21
  • WITHDRAWN — 2

The table

#filefigure(s)claimstatuscommand that reproduces it
1MCP_LISTINGS.md30 26 2026WITHDRAWN: the previous '30 tools' figure, kept as the record of the correctionWITHDRAWNpython claims_audit.py --numbers
2MCP_LISTINGS.md68The MCP server exposes 68 toolsREPRODUCIBLEpython claims_audit.py --numbers
3MCP_LISTINGS.md68The enumerated tool list matches the serverREPRODUCIBLEpython claims_audit.py --numbers
4README.md10 13 3Framework adapters: 10 of 13 verified against current upstream, 3 recorded brokenREPRODUCIBLEpython claims_audit.py --numbers
5README.md0The control: with the guard off we score zero, so the number is the mechanismREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
6README.md0 86.7 13.3 95 3.3 26.7Graphiti keeps the correction 86.7% of the time; resurrection 13.3%, 95% CI [3.3, 26.7]REPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
7README.md2.0.11 53.3 46.7 95 30.0 63.3mem0 2.0.11 keeps the correction 53.3%; resurrection 46.7%, 95% CI [30.0, 63.3]REPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
8README.md30 2.0.11 2026 2.0.18Sample size per system, and the exact competitor version measuredREPRODUCIBLEcurl -s https://pypi.org/pypi/mem0ai/json
9README.md100 0inspeximus keeps a corrected fact 100% of the time; it never resurrects the old valueREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
10README.md30Sample size: 30 trials per systemREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
11README.md13.3The Graphiti row's raw resurrection decomposed: four pre-echo extraction missesREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
12README.md0 26Graphiti's bi-temporal invalidation held 26/26 corrections that were extracted pre-echo; its 13.3% raw resurrection is four extraction misses, not echo failuresREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
13README.md0On echo-attributable resurrection specifically, Graphiti scores 0% -- the separator is whether the supersession link is recorded at write time, not which vendor recorded itREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
14README.md68The MCP server exposes 68 toolsREPRODUCIBLEpython claims_audit.py --numbers
15README.md0Mutation gate: zero seeded defects survivedREPRODUCIBLEpython tools/mutation_check_parallel.py
16README.md98.3 0.01Our own production store: source populated vs actually re-checkableREPRODUCIBLE-WITH-DEPScurl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py
17README.md2,600 175Suite size, and the mutation gate that makes it evidence: 175 seeded, 175 killedREPRODUCIBLEpython tools/mutation_check_parallel.py
18README.md0Zero required dependencies -- every requirement in the wheel is an optional extraREPRODUCIBLEcurl -s https://pypi.org/pypi/inspeximus/json
19claude-code.html68The MCP server exposes 68 toolsREPRODUCIBLEpython claims_audit.py --numbers
20compare.html0The control: with our own guard off, we keep the correction 0% of the timeREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
21compare.html100The objection names our own headline numberREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
22compare.html0 86.7 13.3 95 3.3 26.7Graphiti keeps the correction 86.7%; resurrection 13.3%, 95% CI [3.3, 26.7]REPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
23compare.html100The control restated: guard off resurrects 100% of the timeREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
24compare.html53.3 46.7 95 30.0 63.3mem0 2.0.11 keeps the correction 53.3%; resurrection 46.7%, 95% CI [30.0, 63.3]REPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
25compare.html30Trials per systemREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
26compare.html100 0inspeximus keeps the correction 100% of the time and resurrects the old value 0%REPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
27compare.html30Sample size, stated in the page's own structured data so the two cannot driftREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
28compare.html2026The competitor version and month measured, in the structured dataREPRODUCIBLEcurl -s https://pypi.org/pypi/mem0ai/json
29compare.html2026The month the competitor figure was measured, said in prose next to the claimREPRODUCIBLEcurl -s https://pypi.org/pypi/mem0ai/json
30compare.html30 0Why intervals are shown; and our 0% with the guard onREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
31docs/DEEP_DIVE.md13 0 5The example claims_audit run: 13 checks pass, 5 are not testable from this packageREPRODUCIBLEpython claims_audit.py --local
32docs/DEEP_DIVE.md8The bedrock synthesis was checked from ~8 directionsEXTERNAL
33docs/DEEP_DIVE.md15 18 60 2 9 1 0 8 4regex_extractor chain binding on benchmarks/chain_binding/ (15 chains, 18 unrelated pairs, 60 prose sentences): chains collapsing to one record 2/15 -> 9/15; false binds on unrelated pairs 1/18 -> 0/18; non-declarative prose keyed 8/60 -> 4/60REPRODUCIBLEpython benchmarks/chain_binding/probe.py
34docs/DEEP_DIVE.md0.36Per-memory outcome attribution reaches only ~0.36 power at n-of-1EXTERNAL
35docs/DEEP_DIVE.md66.9 71.2mem0 and Zep's self-reported LLM-judged QA scoresEXTERNAL
36docs/DEEP_DIVE.md4...and by ~4x at one-eighth budgetEXTERNAL
37docs/DEEP_DIVE.md1.8Value-ranked consolidation beats FIFO by ~1.8x at half budgetEXTERNAL
38docs/DEEP_DIVE.md0.17 1.00A soft delete leaves the value recoverable in 5 of 6 stores (0.17); a wired hard delete scores 1.00REPRODUCIBLEpython probes/forget_verification_bench.py
39docs/DEEP_DIVE.md1.00 0.17Same six-store fan-out measurement, restated in the four-operations tableREPRODUCIBLEpython probes/forget_verification_bench.py
40docs/DEEP_DIVE.md0.0000Run-to-run determinism at a fixed instant: arm (a) divergence 0.0000 on every corpusREPRODUCIBLE-WITH-DEPSpython probes/reinforce_accuracy_ablation.py
41docs/DEEP_DIVE.md20Pruning hub notes lifts lexical recall ~20% on a link-spammed store onlyEXTERNAL
42docs/DEEP_DIVE.md0.00 0.05In-repo cross-system echo cell: resurrection rate inspeximus 0.00, mem0 0.05, Graphiti 0.00REPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_echo.py --systems inspeximus
43docs/DEEP_DIVE.md5 0.94 0.25Lexical recall@5 decays 0.94 -> 0.25 as the store growsEXTERNAL
44docs/DEEP_DIVE.md2026 0.78 0.65The superseded pair, quoted inside the note that discharges its caveatREPRODUCIBLE-WITH-DEPSpython benchmarks/locomo/run.py --subset full --retrieval-only
45docs/DEEP_DIVE.md0.83 0.70The copy-paste command with its expected output inlineREPRODUCIBLE-WITH-DEPSpython benchmarks/locomo/run.py --subset full --retrieval-only
46docs/DEEP_DIVE.md0.19 0.29A withdrawn 0.19->0.29 delta, cited as an example of a confound we found and correctedEXTERNAL
47docs/DEEP_DIVE.md1536The LOCOMO question denominator behind the retrieval pairREPRODUCIBLE-WITH-DEPSpython benchmarks/locomo/run.py --subset full --retrieval-only
48docs/DEEP_DIVE.md25 0.83 0.70LOCOMO retrieval-recall@25 = 0.83 (any evidence turn) / 0.70 (all), n=1536, reinforce=FalseREPRODUCIBLE-WITH-DEPSpython benchmarks/locomo/run.py --subset full --retrieval-only
49docs/DEEP_DIVE.md1536,The LoCoMo config size behind recall_any@1PENDING-HARNESSpython probes/retrieval_recall_locomo.py --k 25
50docs/DEEP_DIVE.md0.7839 0.6484 0.783 0.648 1536The OLD published pair reproduces exactly at its own operating point (reinforce=True)REPRODUCIBLE-WITH-DEPSpython benchmarks/locomo/run.py --subset full --retrieval-only
51docs/DEEP_DIVE.md68The MCP server exposes 68 toolsREPRODUCIBLEpython -c "import re,pathlib;print(len(re.findall(chr(64)+chr(109)+chr(99)+chr(112)+chr(46)+'tool', pathlib.Path('inspeximus/mcp_server.py').read_text(encoding='utf-8'))))"
52docs/DEEP_DIVE.md0.592 0.544 2MemOps answer accuracy: keep-all 0.592, mem0 0.544; ~2% of mem0 extractions failed to parseEXTERNAL
53docs/DEEP_DIVE.md0.593MemOps answer accuracy: inspeximus 0.593EXTERNAL
54docs/DEEP_DIVE.md519 917 606 24mem0's default pipeline spends 519-917 s (median 606) of LLM extraction per MemOps scenarioEXTERNAL
55docs/DEEP_DIVE.md2~2% of mem0's MemOps extraction calls failed to parseEXTERNAL
56docs/DEEP_DIVE.md24 50MemOps: 24 long-context scenarios, ~50 sessions eachEXTERNAL
57docs/DEEP_DIVE.md42Operating-point trap: a cosine top-1 store scores 42%REPRODUCIBLE-WITH-DEPSpython probes/operating_point_memory.py
58docs/DEEP_DIVE.md100The layered store scores 100% across all three operating pointsREPRODUCIBLE-WITH-DEPSpython probes/operating_point_memory.py
59docs/DEEP_DIVE.md0 8...and 0/8 on poisonREPRODUCIBLE-WITH-DEPSpython probes/operating_point_memory.py
60docs/DEEP_DIVE.md0 8 67...0/8 on updated facts; a recency store scores 67%REPRODUCIBLE-WITH-DEPSpython probes/operating_point_memory.py
61docs/DEEP_DIVE.md5 0.86 0.20On paraphrase queries semantic recall@5 is 0.86 vs 0.20 lexicalEXTERNAL
62docs/DEEP_DIVE.md0.00 0.57 1.00RAMR ECHO-RESISTANCE: keyed-without-guard 0.00, add-based 0.57, echo_guard 1.00EXTERNAL
63docs/DEEP_DIVE.md0.397recall_any@1 = 0.397 with nomic task prefixes on one LoCoMo configPENDING-HARNESSpython probes/retrieval_recall_locomo.py --k 1
64docs/DEEP_DIVE.md30 2.8At a 30% keep-budget, access-decay retains 2.8% of high-value/low-frequency memoriesEXTERNAL
65docs/DEEP_DIVE.md3 2.2 7~3x more value kept, persisting at ~2.2x even at a 7% budgetEXTERNAL
66docs/DEEP_DIVE.md20 100 64...20% of total value, vs 100% and 64% for the value-aware blendEXTERNAL
67docs/DEEP_DIVE.md0.65 2.6Semantic recall@5 holds ~0.65 at full scale, ~2.6x lexicalEXTERNAL
68docs/DEEP_DIVE.md0.2213NEGATIVE CONTROL: with the salience bar removed, rejection collapses to 0.2213REPRODUCIBLEpython probes/session_digest_multisession.py
69docs/DEEP_DIVE.md7close_session costs 7 ms on the 2,606-record fixtureREPRODUCIBLEpython probes/session_digest_multisession.py
70docs/DEEP_DIVE.md2,606 1.000SessionEnd digest -> SessionStart injection, 8-session / 2,606-record fixture: injection recall 1.000 of a session's conclusions reach the next sessionREPRODUCIBLEpython probes/session_digest_multisession.py
71docs/DEEP_DIVE.md1.0000Below-threshold rejection 1.0000 on the same fixtureREPRODUCIBLEpython probes/session_digest_multisession.py
72docs/DEEP_DIVE.md0 8WITHDRAWN: 'severe-test 8/8' -- the probe reports 0/24 and nothing here produces an 8/8WITHDRAWNpython probes/supersession_replication.py
73docs/DEEP_DIVE.md0.61A cosine classifier separating a contradiction from a rephrase scores AUROC ~0.61REPRODUCIBLE-WITH-DEPSpython probes/supersession_replication.py
74docs/DEEP_DIVE.md0.613 41.7 0.0The 2026-08-01 re-run of that probe, quoted with its dateREPRODUCIBLE-WITH-DEPSpython probes/supersession_replication.py
75docs/DEEP_DIVE.md42A similarity-based store serves the stale value ~42% of the timeREPRODUCIBLE-WITH-DEPSpython probes/supersession_replication.py
76docs/DEEP_DIVE.md0The deterministic SRO key drives the stale-value rate to 0%REPRODUCIBLE-WITH-DEPSpython probes/supersession_replication.py
77docs/DEEP_DIVE.md0.9 10Content-declared corroboration falls to a sybil at ~0.9 attack-success across 10 modelsREPRODUCIBLE-WITH-DEPSpython probes/memory_defense_layer_probe.py
78docs/DEEP_DIVE.md0 80GAP CONTROL: 0 of 80 top-1 answers move when the two reads are not separated at allREPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
79docs/DEEP_DIVE.md80The time-gap measurement runs on four LOCOMO conversations, 80 questions sampled from eachREPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
80docs/DEEP_DIVE.md64 83 320Reading the same untouched store twice ~2s apart moves 64-83 of 320 top-1 answers (four LOCOMO conversations, 80 questions each, reinforce=False)REPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
81docs/DEEP_DIVE.md0.0094Across five randomised insertion orders the hit@1 change over the gap runs +0.0094 to -0.0219REPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
82docs/DEEP_DIVE.md16 80The effect saturates: 16 of 80 move at a two-second gap and the same count at tenREPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
83docs/DEEP_DIVE.md-0.0219 2 5 -0.0062The hit@1 change is negative in only 2 of 5 randomised insert orders; natural conversation order alone reads -0.0062REPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
84docs/DEEP_DIVE.md0.847Every top-1 answer that moved across the gap moved between records reported at the same score, e.g. 0.847 against 0.847REPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
85docs/DEEP_DIVE.md100 6100% of the moved answers stayed inside a displayed tie, across all six insert ordersREPRODUCIBLE-WITH-DEPSpython probes/recall_over_a_time_gap.py
86docs/DEEP_DIVE.md10,000Contradiction detection runs in production over the ~10,000-note vaultEXTERNAL
87docs/DEEP_DIVE.md10,000inspeximus has run daily over a ~10,000-note vaultEXTERNAL
88index.html9 0Homepage counter: 9 framework adaptersREPRODUCIBLEpython -c "import pathlib;print(sorted(p.stem for p in pathlib.Path('inspeximus/integrations').glob('*.py')))"
89index.html0.00Benchmark bar: Graphiti 0.00REPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_revert.py --systems inspeximus,graphiti --n 20
90index.html0.75Benchmark bar: inspeximus 0.75REPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_revert.py --systems inspeximus --n 20
91index.html0.20Benchmark bar: mem0 0.20REPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_revert.py --systems inspeximus,mem0 --n 20
92index.html100The control: with our guard off we resurrect every time, so the number is the mechanismREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
93index.html30Sample size for the native-config echo runREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
94index.html0 13.3 46.7Corrected-fact resurrection per system on their native configsREPRODUCIBLE-WITH-DEPSpython ramr_echo_resistance_backends.py # RAMR repo
95index.html10 13 310 of 13 framework adapters verified against current upstream; 3 recorded brokenREPRODUCIBLEpython tools/integration_conformance.py
96index.html68 0Homepage counter: 68 MCP toolsREPRODUCIBLEpython claims_audit.py --numbers
97index.html68Homepage heading: 68 MCP toolsREPRODUCIBLEpython claims_audit.py --numbers
98index.html2.0.11 2026The exact competitor version and date measured, stated rather than implied as currentREPRODUCIBLEcurl -s https://pypi.org/pypi/mem0ai/json
99index.html0.75 0.20 0.00 20 95Cross-system revert success over n=20: inspeximus 0.75, mem0 0.20, Graphiti 0.00REPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_revert.py --systems inspeximus --n 20
100index.html0.75 0.20 20 0Homepage counter restating the revert cellREPRODUCIBLE-WITH-DEPSpython probes/integrity_bench_revert.py --systems inspeximus --n 20
101index.html98.3Our own store: fraction of records carrying a source fieldREPRODUCIBLE-WITH-DEPScurl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py
102index.html0.01Our own store: fraction whose source actually resolvesREPRODUCIBLE-WITH-DEPScurl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py
103index.html0Homepage counter: 0 runtime dependenciesREPRODUCIBLEpython claims_audit.py --local

Notes

  • mcp-tool-count — Published as 30 until 2026-08-01 -- 26 short -- while the homepage said 15 in one place and 56 in another. Three surfaces, one server, no error anywhere. Now read from the code.
  • readme-adapter-conformance — Read from docs/integration_conformance.json by _live_consistency(), which now checks BOTH index.html and README.md -- a second copy of a number is a second place for it to go stale. The other 12 in this file is the EU AI Act article number and stays a declared non-claim; the counts moved 9/12 -> 10/13 when the llm-errata adapter landed, and this line is why the drift surfaced instead of shipping; COUNT-DRIFT caught the collision the moment this line was added, which is the whole point.
  • readme-echo-graphiti — Measured on the vendor's own native config (Neo4j + OpenAI), n=30.
  • readme-echo-mem0 — Version-stamped on purpose: mem0 is on 2.0.18 as of 2026-08-11 and we have NOT re-run it.
  • readme-mcp-tool-count — Checked against the live @mcp.tool() count by _live_consistency(), not read from here.
  • readme-own-source-coverage — Published as our own failure, not a product claim. 210,499 records across ten stores.
  • cc-tool-count — Checked against the live @mcp.tool() count by _live_consistency(), not read from here.
  • cmp-graphiti — The bare 0 on this line is the MAJOR VERSION in "Graphiti 0.x", not a measurement. It is listed rather than excused, because declaring "0" a non-claim file-wide would also excuse the two real zeros this page publishes.
  • cmp-mem0 — Version-stamped deliberately: mem0 is on 2.0.18 and we have NOT re-run it.
  • readme-audit-summary — Self-referential, so it is checked against len(CHECKS) and len(NOT_TESTABLE_HERE) rather than trusted. The block used to name inspeximus-1.24.1 while the package was at 1.89.0; the version line was dropped rather than pinned, because it would go stale on every release.
  • readme-bedrock-directions — A count of the analytical directions taken, not a measurement. Left in because the sentence labels itself 'a synthesis over those cases, not a proof'.
  • readme-chain-binding — The 'before' column is measured against git show main:inspeximus/core.py on the same fixture, not quoted from elsewhere. The false-bind row is the control: a keyer that binds everything scores a perfect 15/15 while tripping all 18 negative pairs, which is why the bind rate alone is not evidence.
  • readme-competitor-judges — Other projects' published numbers, cited as not comparable across harnesses -- which is the point the sentence makes.
  • readme-erasure-fanout-table — Replaced a 'measured 15/15 on a verified-forgetting severe-test' for which no artifact in this repository produces a 15/15 of anything. The bench that DOES exist scores 0.17 / 1.00 over six stores, so the sentence now cites the number the committed code prints.
  • readme-integrity-echo-cell — The inspeximus column runs locally and free; the mem0/Graphiti columns need OPENAI_API_KEY and a live neo4j, which is why this is WITH-DEPS rather than REPRODUCIBLE.
  • readme-lexical-decay — Agora Lab b4c260, cited in place. No probe in this repository reproduces it.
  • readme-locomo-confound — Kept deliberately: it is a retraction, not a claim. Removing it would erase the correction.
  • readme-locomo-headline — This row was PENDING-HARNESS for the whole of this audit, and it is the reason that status exists. The harness landed as benchmarks/locomo/ and the row moved WITHOUT the number having been re-asserted in the meantime. Verified here against the committed result benchmarks/locomo/results/full_retrieval.json rather than against the prose: recall_any 0.8262 / recall_all 0.6986 on the published 1536-question denominator, pinned at k=25, mode=hybrid, prefer=speaker, reinforce=false. WITH-DEPS because LOCOMO is not ours to redistribute -- the command needs locomo10.json downloaded (sha256 pinned in config.json), though the committed result is readable without it.
  • readme-locomo-old-pair — The pair this audit opened on. It was never wrong -- it was measured with recall()'s reinforce=True default, which mutates value/last_access, so each benchmark query was answered by a store the previous queries had modified and the score depended on question order. Pinning reinforce=False makes the run deterministic and scores 4-5 points HIGHER. A number that moves when you fix the instrument is exactly what an unreproducible number hides.
  • readme-mcp-tools — Checked against the live @mcp.tool() count by _live_consistency(), not by reading it here.
  • readme-memops-cost — Agora MemOps harness; needs mem0 + an LLM budget, so it cannot ship here.
  • readme-memops-scenarios — Harness lives in the Agora repo (agora_output/lab/memops), already linked in place.
  • readme-operating-cosine — Needs a local nomic-embed-text (Ollama).
  • readme-paraphrase — Agora Lab 3501f1.
  • readme-ramr-echo — From RAMR, a separate repository. The README presented these as if produced here; it now says where they come from AND points at this repo's own echo cell, which measures a different quantity and does NOT flatter us.
  • readme-recall-any1 — Same dataset blocker as the headline pair; already flagged in place as not reproducible here.
  • readme-retention-cold — Agora Lab 19d802.
  • readme-semantic-hold — Agora Lab b4c260.
  • readme-session-digest-control — Registered deliberately rather than dropped: without it a rejection of 1.0000 cannot be told apart from a fixture that contained nothing to reject.
  • readme-supersession-auroc — Re-run 2026-08-01: AUROC 0.613. Needs a local nomic-embed-text (Ollama) and numpy.
  • readme-supersession-stale — Re-run 2026-08-01: 41.7%.
  • readme-sybil-attack — The harness is committed; reproducing the number needs ten models and a judge, which no checkout can ship.
  • readme-time-gap-cause — Registered deliberately rather than dropped: without a zero-gap arm, 'the ranking depends on when you ask' is only an observation about two reads and names no cause.
  • readme-time-gap-spread-low — The natural-order figure is registered beside the spread on purpose: alone it reproduces to four decimals every run and reads as a systematic loss, which is the fixture (LOCOMO gold turns skew late, so gold records are newer) and not a property of the store.
  • readme-vault-contradictions — Same private deployment as the hero line.
  • readme-vault-hero — Our own private Obsidian vault. There is no command; the text now says so instead of implying the reader could check it.
  • site-adapters — Was 6 while the README said nine and the package ships nine agent-framework adapters (autogen, crewai, google_adk, haystack, langchain, langgraph, llamaindex, openai_agents, pydantic_ai).
  • site-echo-row — Replaced a STALE caveat claiming all three tie on this cell; the later run separates them.
  • site-integration-conformance — Read from the committed ledger docs/integration_conformance.json by _live_consistency(), not typed. The page previously said 'Drop-in for' all nine frameworks with no qualifier at all, while crewai 1.15.6, openai-agents 0.18.3 and langgraph-checkpointer 1.2.9 were recorded broken -- an unqualified capability claim contradicted by a JSON file in the same repo.
  • site-mcp-tools-counter — Was 15. The counter renders data-count, so the figure a reader sees lives in an attribute -- which is why the scanner hoists data-count out of the tag before stripping tags.
  • site-revert-bench — The inspeximus column runs locally; mem0 needs OPENAI_API_KEY and Graphiti a live neo4j. Methodology and CIs: probes/INTEGRITY_BENCHMARK.md.
  • site-zero-deps — This is the c_zero_deps check, which reads installed METADATA or, failing that, the declared pyproject dependencies -- and hard-fails if it can read neither.

Known unenforced numbers

These are outside the 6 token-enforced files, so the guard above does not cover them. They are listed because "absent from the table" and "not a problem" are different statements, and here only the first one is true. None may be promoted onto the reader-facing surface while it still says PENDING-HARNESS.

wherefigurestatuswhy it is here
inspeximus/core.py — recall_iterative docstring0.057 -> 0.186 (3.3x)PENDING-HARNESSQuoted with no scope. That ratio is n=70 over the THREE HARDEST LoCoMo conversations; across the full benchmark it is 0.145 -> 0.297 (2.05x), n=276, all ten conversations. A flattering subset ratio published without the subset is the exact defect this audit exists to find, and neither figure reproduces from a clean checkout (same LOCOMO dataset blocker as readme-locomo-headline). Scope it to the subset AND give the full-benchmark pair, or drop it, once a harness lands.
docs / docstrings — the recall tie policyn/a (a stated policy, not a figure)PENDING-HARNESS'equal relevance => newest first' is measured to FAIL in hybrid and auto modes: RRF gives equivalent records distinct fused scores, so the tie-break never fires and top-1 is the OLDEST. Anywhere the policy is asserted unconditionally it needs scoping to the modes where it holds.
bench/ — MemoryAgentBench Conflict Resolutionany CR scorePENDING-HARNESSOur own gate killed the 'supersession wins CR' headline: a naive keep-all store ties us (97 vs 96, and 85 vs 87 on the faithful re-run), and a 6k-vs-32k context discrepancy in bench/README.md is still open. No CR number may go onto the reader-facing surface until that is resolved.
index.html / MCP_LISTINGS.md / README.md — the MCP tool count60REPRODUCIBLENo longer hypothetical. Between this audit and its rebase the server grew from 56 to 58 tools; a sibling corrected ONE of the four places that publish the count (the homepage heading) and left the homepage counter, both MCP_LISTINGS figures and the README at 56. This audit named all four on its first run after the rebase. Checked against the live @mcp.tool() count, never typed.
bench/README.md — its own committed JSON9 of 12 cellsPENDING-HARNESSReported to disagree with the JSON it is generated from. Outside this audit's token-enforced scope (README.md, MCP_LISTINGS.md, index.html), and left to the unit that owns the reconciliation rather than guessed at from here.
docs / homepage — LongMemEval end-to-end0.45, vs a 0.50 oracle ceiling and a 0.05 no-memory floor, n=20PENDING-HARNESSLanded as a pilot. No figure from it has reached README.md, MCP_LISTINGS.md or index.html, and none should until it carries its scope -- n=20, a pilot, and a band check that exits 5 by design.

What was withdrawn, and why that is the point

Removing a number we cannot back is a win. Every WITHDRAWN row above is a figure this audit deleted from the reader-facing surface rather than dress up: either no artifact in this repository produces it, or the artifact it named does not exist. Showing the gap is what makes the rest worth reading.