Published numbers
August 13, 2026 · View on GitHub
A number a reader cannot reproduce is not evidence. This table is the whole reader-facing surface's quantitative content, each figure against the exact command that produces it.
Scope, stated so a green result cannot be over-read. The audit enforces registration on
README.md, MCP_LISTINGS.md and index.html: every numeric token a reader sees on those three
must be a row below or a declared non-claim, and python claims_audit.py --numbers fails otherwise.
CHANGELOG.md and docs/ are not token-enforced — they are covered only by the weaker check
that every artifact path they name exists. Their numbers are not audited here, and reading this
page as "every number in the project is backed" would be exactly the over-read it exists to prevent.
The ratio
- 350 numeric tokens are published across the 6 enforced files: README.md, docs/DEEP_DIVE.md, MCP_LISTINGS.md, index.html, compare.html, claude-code.html.
- 190 of those are quantitative claims, in 103 registry rows below.
- 78 rows (78/103) are reproducible by a command committed to this repository
(
REPRODUCIBLEneeds nothing but this checkout;REPRODUCIBLE-WITH-DEPSneeds a service or dataset we cannot redistribute, named in the command column). - The remaining 25 are
PENDING-HARNESS,EXTERNALorWITHDRAWN. - The other 160 tokens are declared non-claims — citation years, article numbers, ordinals, ports, example literals — each with a reason and an exact expected count, so adding one silently is not possible either.
Counts by status:
REPRODUCIBLE— 26REPRODUCIBLE-WITH-DEPS— 52PENDING-HARNESS— 2EXTERNAL— 21WITHDRAWN— 2
The table
| # | file | figure(s) | claim | status | command that reproduces it |
|---|---|---|---|---|---|
| 1 | MCP_LISTINGS.md | 30 26 2026 | WITHDRAWN: the previous '30 tools' figure, kept as the record of the correction | WITHDRAWN | python claims_audit.py --numbers |
| 2 | MCP_LISTINGS.md | 68 | The MCP server exposes 68 tools | REPRODUCIBLE | python claims_audit.py --numbers |
| 3 | MCP_LISTINGS.md | 68 | The enumerated tool list matches the server | REPRODUCIBLE | python claims_audit.py --numbers |
| 4 | README.md | 10 13 3 | Framework adapters: 10 of 13 verified against current upstream, 3 recorded broken | REPRODUCIBLE | python claims_audit.py --numbers |
| 5 | README.md | 0 | The control: with the guard off we score zero, so the number is the mechanism | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 6 | README.md | 0 86.7 13.3 95 3.3 26.7 | Graphiti keeps the correction 86.7% of the time; resurrection 13.3%, 95% CI [3.3, 26.7] | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 7 | README.md | 2.0.11 53.3 46.7 95 30.0 63.3 | mem0 2.0.11 keeps the correction 53.3%; resurrection 46.7%, 95% CI [30.0, 63.3] | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 8 | README.md | 30 2.0.11 2026 2.0.18 | Sample size per system, and the exact competitor version measured | REPRODUCIBLE | curl -s https://pypi.org/pypi/mem0ai/json |
| 9 | README.md | 100 0 | inspeximus keeps a corrected fact 100% of the time; it never resurrects the old value | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 10 | README.md | 30 | Sample size: 30 trials per system | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 11 | README.md | 13.3 | The Graphiti row's raw resurrection decomposed: four pre-echo extraction misses | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 12 | README.md | 0 26 | Graphiti's bi-temporal invalidation held 26/26 corrections that were extracted pre-echo; its 13.3% raw resurrection is four extraction misses, not echo failures | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 13 | README.md | 0 | On echo-attributable resurrection specifically, Graphiti scores 0% -- the separator is whether the supersession link is recorded at write time, not which vendor recorded it | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 14 | README.md | 68 | The MCP server exposes 68 tools | REPRODUCIBLE | python claims_audit.py --numbers |
| 15 | README.md | 0 | Mutation gate: zero seeded defects survived | REPRODUCIBLE | python tools/mutation_check_parallel.py |
| 16 | README.md | 98.3 0.01 | Our own production store: source populated vs actually re-checkable | REPRODUCIBLE-WITH-DEPS | curl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py |
| 17 | README.md | 2,600 175 | Suite size, and the mutation gate that makes it evidence: 175 seeded, 175 killed | REPRODUCIBLE | python tools/mutation_check_parallel.py |
| 18 | README.md | 0 | Zero required dependencies -- every requirement in the wheel is an optional extra | REPRODUCIBLE | curl -s https://pypi.org/pypi/inspeximus/json |
| 19 | claude-code.html | 68 | The MCP server exposes 68 tools | REPRODUCIBLE | python claims_audit.py --numbers |
| 20 | compare.html | 0 | The control: with our own guard off, we keep the correction 0% of the time | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 21 | compare.html | 100 | The objection names our own headline number | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 22 | compare.html | 0 86.7 13.3 95 3.3 26.7 | Graphiti keeps the correction 86.7%; resurrection 13.3%, 95% CI [3.3, 26.7] | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 23 | compare.html | 100 | The control restated: guard off resurrects 100% of the time | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 24 | compare.html | 53.3 46.7 95 30.0 63.3 | mem0 2.0.11 keeps the correction 53.3%; resurrection 46.7%, 95% CI [30.0, 63.3] | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 25 | compare.html | 30 | Trials per system | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 26 | compare.html | 100 0 | inspeximus keeps the correction 100% of the time and resurrects the old value 0% | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 27 | compare.html | 30 | Sample size, stated in the page's own structured data so the two cannot drift | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 28 | compare.html | 2026 | The competitor version and month measured, in the structured data | REPRODUCIBLE | curl -s https://pypi.org/pypi/mem0ai/json |
| 29 | compare.html | 2026 | The month the competitor figure was measured, said in prose next to the claim | REPRODUCIBLE | curl -s https://pypi.org/pypi/mem0ai/json |
| 30 | compare.html | 30 0 | Why intervals are shown; and our 0% with the guard on | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 31 | docs/DEEP_DIVE.md | 13 0 5 | The example claims_audit run: 13 checks pass, 5 are not testable from this package | REPRODUCIBLE | python claims_audit.py --local |
| 32 | docs/DEEP_DIVE.md | 8 | The bedrock synthesis was checked from ~8 directions | EXTERNAL | — |
| 33 | docs/DEEP_DIVE.md | 15 18 60 2 9 1 0 8 4 | regex_extractor chain binding on benchmarks/chain_binding/ (15 chains, 18 unrelated pairs, 60 prose sentences): chains collapsing to one record 2/15 -> 9/15; false binds on unrelated pairs 1/18 -> 0/18; non-declarative prose keyed 8/60 -> 4/60 | REPRODUCIBLE | python benchmarks/chain_binding/probe.py |
| 34 | docs/DEEP_DIVE.md | 0.36 | Per-memory outcome attribution reaches only ~0.36 power at n-of-1 | EXTERNAL | — |
| 35 | docs/DEEP_DIVE.md | 66.9 71.2 | mem0 and Zep's self-reported LLM-judged QA scores | EXTERNAL | — |
| 36 | docs/DEEP_DIVE.md | 4 | ...and by ~4x at one-eighth budget | EXTERNAL | — |
| 37 | docs/DEEP_DIVE.md | 1.8 | Value-ranked consolidation beats FIFO by ~1.8x at half budget | EXTERNAL | — |
| 38 | docs/DEEP_DIVE.md | 0.17 1.00 | A soft delete leaves the value recoverable in 5 of 6 stores (0.17); a wired hard delete scores 1.00 | REPRODUCIBLE | python probes/forget_verification_bench.py |
| 39 | docs/DEEP_DIVE.md | 1.00 0.17 | Same six-store fan-out measurement, restated in the four-operations table | REPRODUCIBLE | python probes/forget_verification_bench.py |
| 40 | docs/DEEP_DIVE.md | 0.0000 | Run-to-run determinism at a fixed instant: arm (a) divergence 0.0000 on every corpus | REPRODUCIBLE-WITH-DEPS | python probes/reinforce_accuracy_ablation.py |
| 41 | docs/DEEP_DIVE.md | 20 | Pruning hub notes lifts lexical recall ~20% on a link-spammed store only | EXTERNAL | — |
| 42 | docs/DEEP_DIVE.md | 0.00 0.05 | In-repo cross-system echo cell: resurrection rate inspeximus 0.00, mem0 0.05, Graphiti 0.00 | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_echo.py --systems inspeximus |
| 43 | docs/DEEP_DIVE.md | 5 0.94 0.25 | Lexical recall@5 decays 0.94 -> 0.25 as the store grows | EXTERNAL | — |
| 44 | docs/DEEP_DIVE.md | 2026 0.78 0.65 | The superseded pair, quoted inside the note that discharges its caveat | REPRODUCIBLE-WITH-DEPS | python benchmarks/locomo/run.py --subset full --retrieval-only |
| 45 | docs/DEEP_DIVE.md | 0.83 0.70 | The copy-paste command with its expected output inline | REPRODUCIBLE-WITH-DEPS | python benchmarks/locomo/run.py --subset full --retrieval-only |
| 46 | docs/DEEP_DIVE.md | 0.19 0.29 | A withdrawn 0.19->0.29 delta, cited as an example of a confound we found and corrected | EXTERNAL | — |
| 47 | docs/DEEP_DIVE.md | 1536 | The LOCOMO question denominator behind the retrieval pair | REPRODUCIBLE-WITH-DEPS | python benchmarks/locomo/run.py --subset full --retrieval-only |
| 48 | docs/DEEP_DIVE.md | 25 0.83 0.70 | LOCOMO retrieval-recall@25 = 0.83 (any evidence turn) / 0.70 (all), n=1536, reinforce=False | REPRODUCIBLE-WITH-DEPS | python benchmarks/locomo/run.py --subset full --retrieval-only |
| 49 | docs/DEEP_DIVE.md | 1536, | The LoCoMo config size behind recall_any@1 | PENDING-HARNESS | python probes/retrieval_recall_locomo.py --k 25 |
| 50 | docs/DEEP_DIVE.md | 0.7839 0.6484 0.783 0.648 1536 | The OLD published pair reproduces exactly at its own operating point (reinforce=True) | REPRODUCIBLE-WITH-DEPS | python benchmarks/locomo/run.py --subset full --retrieval-only |
| 51 | docs/DEEP_DIVE.md | 68 | The MCP server exposes 68 tools | REPRODUCIBLE | python -c "import re,pathlib;print(len(re.findall(chr(64)+chr(109)+chr(99)+chr(112)+chr(46)+'tool', pathlib.Path('inspeximus/mcp_server.py').read_text(encoding='utf-8'))))" |
| 52 | docs/DEEP_DIVE.md | 0.592 0.544 2 | MemOps answer accuracy: keep-all 0.592, mem0 0.544; ~2% of mem0 extractions failed to parse | EXTERNAL | — |
| 53 | docs/DEEP_DIVE.md | 0.593 | MemOps answer accuracy: inspeximus 0.593 | EXTERNAL | — |
| 54 | docs/DEEP_DIVE.md | 519 917 606 24 | mem0's default pipeline spends 519-917 s (median 606) of LLM extraction per MemOps scenario | EXTERNAL | — |
| 55 | docs/DEEP_DIVE.md | 2 | ~2% of mem0's MemOps extraction calls failed to parse | EXTERNAL | — |
| 56 | docs/DEEP_DIVE.md | 24 50 | MemOps: 24 long-context scenarios, ~50 sessions each | EXTERNAL | — |
| 57 | docs/DEEP_DIVE.md | 42 | Operating-point trap: a cosine top-1 store scores 42% | REPRODUCIBLE-WITH-DEPS | python probes/operating_point_memory.py |
| 58 | docs/DEEP_DIVE.md | 100 | The layered store scores 100% across all three operating points | REPRODUCIBLE-WITH-DEPS | python probes/operating_point_memory.py |
| 59 | docs/DEEP_DIVE.md | 0 8 | ...and 0/8 on poison | REPRODUCIBLE-WITH-DEPS | python probes/operating_point_memory.py |
| 60 | docs/DEEP_DIVE.md | 0 8 67 | ...0/8 on updated facts; a recency store scores 67% | REPRODUCIBLE-WITH-DEPS | python probes/operating_point_memory.py |
| 61 | docs/DEEP_DIVE.md | 5 0.86 0.20 | On paraphrase queries semantic recall@5 is 0.86 vs 0.20 lexical | EXTERNAL | — |
| 62 | docs/DEEP_DIVE.md | 0.00 0.57 1.00 | RAMR ECHO-RESISTANCE: keyed-without-guard 0.00, add-based 0.57, echo_guard 1.00 | EXTERNAL | — |
| 63 | docs/DEEP_DIVE.md | 0.397 | recall_any@1 = 0.397 with nomic task prefixes on one LoCoMo config | PENDING-HARNESS | python probes/retrieval_recall_locomo.py --k 1 |
| 64 | docs/DEEP_DIVE.md | 30 2.8 | At a 30% keep-budget, access-decay retains 2.8% of high-value/low-frequency memories | EXTERNAL | — |
| 65 | docs/DEEP_DIVE.md | 3 2.2 7 | ~3x more value kept, persisting at ~2.2x even at a 7% budget | EXTERNAL | — |
| 66 | docs/DEEP_DIVE.md | 20 100 64 | ...20% of total value, vs 100% and 64% for the value-aware blend | EXTERNAL | — |
| 67 | docs/DEEP_DIVE.md | 0.65 2.6 | Semantic recall@5 holds ~0.65 at full scale, ~2.6x lexical | EXTERNAL | — |
| 68 | docs/DEEP_DIVE.md | 0.2213 | NEGATIVE CONTROL: with the salience bar removed, rejection collapses to 0.2213 | REPRODUCIBLE | python probes/session_digest_multisession.py |
| 69 | docs/DEEP_DIVE.md | 7 | close_session costs 7 ms on the 2,606-record fixture | REPRODUCIBLE | python probes/session_digest_multisession.py |
| 70 | docs/DEEP_DIVE.md | 2,606 1.000 | SessionEnd digest -> SessionStart injection, 8-session / 2,606-record fixture: injection recall 1.000 of a session's conclusions reach the next session | REPRODUCIBLE | python probes/session_digest_multisession.py |
| 71 | docs/DEEP_DIVE.md | 1.0000 | Below-threshold rejection 1.0000 on the same fixture | REPRODUCIBLE | python probes/session_digest_multisession.py |
| 72 | docs/DEEP_DIVE.md | 0 8 | WITHDRAWN: 'severe-test 8/8' -- the probe reports 0/24 and nothing here produces an 8/8 | WITHDRAWN | python probes/supersession_replication.py |
| 73 | docs/DEEP_DIVE.md | 0.61 | A cosine classifier separating a contradiction from a rephrase scores AUROC ~0.61 | REPRODUCIBLE-WITH-DEPS | python probes/supersession_replication.py |
| 74 | docs/DEEP_DIVE.md | 0.613 41.7 0.0 | The 2026-08-01 re-run of that probe, quoted with its date | REPRODUCIBLE-WITH-DEPS | python probes/supersession_replication.py |
| 75 | docs/DEEP_DIVE.md | 42 | A similarity-based store serves the stale value ~42% of the time | REPRODUCIBLE-WITH-DEPS | python probes/supersession_replication.py |
| 76 | docs/DEEP_DIVE.md | 0 | The deterministic SRO key drives the stale-value rate to 0% | REPRODUCIBLE-WITH-DEPS | python probes/supersession_replication.py |
| 77 | docs/DEEP_DIVE.md | 0.9 10 | Content-declared corroboration falls to a sybil at ~0.9 attack-success across 10 models | REPRODUCIBLE-WITH-DEPS | python probes/memory_defense_layer_probe.py |
| 78 | docs/DEEP_DIVE.md | 0 80 | GAP CONTROL: 0 of 80 top-1 answers move when the two reads are not separated at all | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 79 | docs/DEEP_DIVE.md | 80 | The time-gap measurement runs on four LOCOMO conversations, 80 questions sampled from each | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 80 | docs/DEEP_DIVE.md | 64 83 320 | Reading the same untouched store twice ~2s apart moves 64-83 of 320 top-1 answers (four LOCOMO conversations, 80 questions each, reinforce=False) | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 81 | docs/DEEP_DIVE.md | 0.0094 | Across five randomised insertion orders the hit@1 change over the gap runs +0.0094 to -0.0219 | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 82 | docs/DEEP_DIVE.md | 16 80 | The effect saturates: 16 of 80 move at a two-second gap and the same count at ten | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 83 | docs/DEEP_DIVE.md | -0.0219 2 5 -0.0062 | The hit@1 change is negative in only 2 of 5 randomised insert orders; natural conversation order alone reads -0.0062 | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 84 | docs/DEEP_DIVE.md | 0.847 | Every top-1 answer that moved across the gap moved between records reported at the same score, e.g. 0.847 against 0.847 | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 85 | docs/DEEP_DIVE.md | 100 6 | 100% of the moved answers stayed inside a displayed tie, across all six insert orders | REPRODUCIBLE-WITH-DEPS | python probes/recall_over_a_time_gap.py |
| 86 | docs/DEEP_DIVE.md | 10,000 | Contradiction detection runs in production over the ~10,000-note vault | EXTERNAL | — |
| 87 | docs/DEEP_DIVE.md | 10,000 | inspeximus has run daily over a ~10,000-note vault | EXTERNAL | — |
| 88 | index.html | 9 0 | Homepage counter: 9 framework adapters | REPRODUCIBLE | python -c "import pathlib;print(sorted(p.stem for p in pathlib.Path('inspeximus/integrations').glob('*.py')))" |
| 89 | index.html | 0.00 | Benchmark bar: Graphiti 0.00 | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_revert.py --systems inspeximus,graphiti --n 20 |
| 90 | index.html | 0.75 | Benchmark bar: inspeximus 0.75 | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_revert.py --systems inspeximus --n 20 |
| 91 | index.html | 0.20 | Benchmark bar: mem0 0.20 | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_revert.py --systems inspeximus,mem0 --n 20 |
| 92 | index.html | 100 | The control: with our guard off we resurrect every time, so the number is the mechanism | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 93 | index.html | 30 | Sample size for the native-config echo run | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 94 | index.html | 0 13.3 46.7 | Corrected-fact resurrection per system on their native configs | REPRODUCIBLE-WITH-DEPS | python ramr_echo_resistance_backends.py # RAMR repo |
| 95 | index.html | 10 13 3 | 10 of 13 framework adapters verified against current upstream; 3 recorded broken | REPRODUCIBLE | python tools/integration_conformance.py |
| 96 | index.html | 68 0 | Homepage counter: 68 MCP tools | REPRODUCIBLE | python claims_audit.py --numbers |
| 97 | index.html | 68 | Homepage heading: 68 MCP tools | REPRODUCIBLE | python claims_audit.py --numbers |
| 98 | index.html | 2.0.11 2026 | The exact competitor version and date measured, stated rather than implied as current | REPRODUCIBLE | curl -s https://pypi.org/pypi/mem0ai/json |
| 99 | index.html | 0.75 0.20 0.00 20 95 | Cross-system revert success over n=20: inspeximus 0.75, mem0 0.20, Graphiti 0.00 | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_revert.py --systems inspeximus --n 20 |
| 100 | index.html | 0.75 0.20 20 0 | Homepage counter restating the revert cell | REPRODUCIBLE-WITH-DEPS | python probes/integrity_bench_revert.py --systems inspeximus --n 20 |
| 101 | index.html | 98.3 | Our own store: fraction of records carrying a source field | REPRODUCIBLE-WITH-DEPS | curl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py |
| 102 | index.html | 0.01 | Our own store: fraction whose source actually resolves | REPRODUCIBLE-WITH-DEPS | curl -sO https://raw.githubusercontent.com/DanceNitra/agora/main/research/probes/can_we_reconcile_our_own_index.py && python can_we_reconcile_our_own_index.py |
| 103 | index.html | 0 | Homepage counter: 0 runtime dependencies | REPRODUCIBLE | python claims_audit.py --local |
Notes
- mcp-tool-count — Published as 30 until 2026-08-01 -- 26 short -- while the homepage said 15 in one place and 56 in another. Three surfaces, one server, no error anywhere. Now read from the code.
- readme-adapter-conformance — Read from docs/integration_conformance.json by _live_consistency(), which now checks BOTH index.html and README.md -- a second copy of a number is a second place for it to go stale. The other 12 in this file is the EU AI Act article number and stays a declared non-claim; the counts moved 9/12 -> 10/13 when the llm-errata adapter landed, and this line is why the drift surfaced instead of shipping; COUNT-DRIFT caught the collision the moment this line was added, which is the whole point.
- readme-echo-graphiti — Measured on the vendor's own native config (Neo4j + OpenAI), n=30.
- readme-echo-mem0 — Version-stamped on purpose: mem0 is on 2.0.18 as of 2026-08-11 and we have NOT re-run it.
- readme-mcp-tool-count — Checked against the live @mcp.tool() count by _live_consistency(), not read from here.
- readme-own-source-coverage — Published as our own failure, not a product claim. 210,499 records across ten stores.
- cc-tool-count — Checked against the live @mcp.tool() count by _live_consistency(), not read from here.
- cmp-graphiti — The bare 0 on this line is the MAJOR VERSION in "Graphiti 0.x", not a measurement. It is listed rather than excused, because declaring "0" a non-claim file-wide would also excuse the two real zeros this page publishes.
- cmp-mem0 — Version-stamped deliberately: mem0 is on 2.0.18 and we have NOT re-run it.
- readme-audit-summary — Self-referential, so it is checked against len(CHECKS) and len(NOT_TESTABLE_HERE) rather than trusted. The block used to name inspeximus-1.24.1 while the package was at 1.89.0; the version line was dropped rather than pinned, because it would go stale on every release.
- readme-bedrock-directions — A count of the analytical directions taken, not a measurement. Left in because the sentence labels itself 'a synthesis over those cases, not a proof'.
- readme-chain-binding — The 'before' column is measured against
git show main:inspeximus/core.pyon the same fixture, not quoted from elsewhere. The false-bind row is the control: a keyer that binds everything scores a perfect 15/15 while tripping all 18 negative pairs, which is why the bind rate alone is not evidence. - readme-competitor-judges — Other projects' published numbers, cited as not comparable across harnesses -- which is the point the sentence makes.
- readme-erasure-fanout-table — Replaced a 'measured 15/15 on a verified-forgetting severe-test' for which no artifact in this repository produces a 15/15 of anything. The bench that DOES exist scores 0.17 / 1.00 over six stores, so the sentence now cites the number the committed code prints.
- readme-integrity-echo-cell — The inspeximus column runs locally and free; the mem0/Graphiti columns need OPENAI_API_KEY and a live neo4j, which is why this is WITH-DEPS rather than REPRODUCIBLE.
- readme-lexical-decay — Agora Lab b4c260, cited in place. No probe in this repository reproduces it.
- readme-locomo-confound — Kept deliberately: it is a retraction, not a claim. Removing it would erase the correction.
- readme-locomo-headline — This row was PENDING-HARNESS for the whole of this audit, and it is the reason that status exists. The harness landed as benchmarks/locomo/ and the row moved WITHOUT the number having been re-asserted in the meantime. Verified here against the committed result benchmarks/locomo/results/full_retrieval.json rather than against the prose: recall_any 0.8262 / recall_all 0.6986 on the published 1536-question denominator, pinned at k=25, mode=hybrid, prefer=speaker, reinforce=false. WITH-DEPS because LOCOMO is not ours to redistribute -- the command needs locomo10.json downloaded (sha256 pinned in config.json), though the committed result is readable without it.
- readme-locomo-old-pair — The pair this audit opened on. It was never wrong -- it was measured with recall()'s reinforce=True default, which mutates value/last_access, so each benchmark query was answered by a store the previous queries had modified and the score depended on question order. Pinning reinforce=False makes the run deterministic and scores 4-5 points HIGHER. A number that moves when you fix the instrument is exactly what an unreproducible number hides.
- readme-mcp-tools — Checked against the live @mcp.tool() count by _live_consistency(), not by reading it here.
- readme-memops-cost — Agora MemOps harness; needs mem0 + an LLM budget, so it cannot ship here.
- readme-memops-scenarios — Harness lives in the Agora repo (agora_output/lab/memops), already linked in place.
- readme-operating-cosine — Needs a local nomic-embed-text (Ollama).
- readme-paraphrase — Agora Lab 3501f1.
- readme-ramr-echo — From RAMR, a separate repository. The README presented these as if produced here; it now says where they come from AND points at this repo's own echo cell, which measures a different quantity and does NOT flatter us.
- readme-recall-any1 — Same dataset blocker as the headline pair; already flagged in place as not reproducible here.
- readme-retention-cold — Agora Lab 19d802.
- readme-semantic-hold — Agora Lab b4c260.
- readme-session-digest-control — Registered deliberately rather than dropped: without it a rejection of 1.0000 cannot be told apart from a fixture that contained nothing to reject.
- readme-supersession-auroc — Re-run 2026-08-01: AUROC 0.613. Needs a local nomic-embed-text (Ollama) and numpy.
- readme-supersession-stale — Re-run 2026-08-01: 41.7%.
- readme-sybil-attack — The harness is committed; reproducing the number needs ten models and a judge, which no checkout can ship.
- readme-time-gap-cause — Registered deliberately rather than dropped: without a zero-gap arm, 'the ranking depends on when you ask' is only an observation about two reads and names no cause.
- readme-time-gap-spread-low — The natural-order figure is registered beside the spread on purpose: alone it reproduces to four decimals every run and reads as a systematic loss, which is the fixture (LOCOMO gold turns skew late, so gold records are newer) and not a property of the store.
- readme-vault-contradictions — Same private deployment as the hero line.
- readme-vault-hero — Our own private Obsidian vault. There is no command; the text now says so instead of implying the reader could check it.
- site-adapters — Was 6 while the README said nine and the package ships nine agent-framework adapters (autogen, crewai, google_adk, haystack, langchain, langgraph, llamaindex, openai_agents, pydantic_ai).
- site-echo-row — Replaced a STALE caveat claiming all three tie on this cell; the later run separates them.
- site-integration-conformance — Read from the committed ledger docs/integration_conformance.json by _live_consistency(), not typed. The page previously said 'Drop-in for' all nine frameworks with no qualifier at all, while crewai 1.15.6, openai-agents 0.18.3 and langgraph-checkpointer 1.2.9 were recorded broken -- an unqualified capability claim contradicted by a JSON file in the same repo.
- site-mcp-tools-counter — Was 15. The counter renders data-count, so the figure a reader sees lives in an attribute -- which is why the scanner hoists data-count out of the tag before stripping tags.
- site-revert-bench — The inspeximus column runs locally; mem0 needs OPENAI_API_KEY and Graphiti a live neo4j. Methodology and CIs: probes/INTEGRITY_BENCHMARK.md.
- site-zero-deps — This is the c_zero_deps check, which reads installed METADATA or, failing that, the declared pyproject dependencies -- and hard-fails if it can read neither.
Known unenforced numbers
These are outside the 6 token-enforced files, so the guard above does not cover them.
They are listed because "absent from the table" and "not a problem" are different statements,
and here only the first one is true. None may be promoted onto the reader-facing surface while it
still says PENDING-HARNESS.
| where | figure | status | why it is here |
|---|---|---|---|
inspeximus/core.py — recall_iterative docstring | 0.057 -> 0.186 (3.3x) | PENDING-HARNESS | Quoted with no scope. That ratio is n=70 over the THREE HARDEST LoCoMo conversations; across the full benchmark it is 0.145 -> 0.297 (2.05x), n=276, all ten conversations. A flattering subset ratio published without the subset is the exact defect this audit exists to find, and neither figure reproduces from a clean checkout (same LOCOMO dataset blocker as readme-locomo-headline). Scope it to the subset AND give the full-benchmark pair, or drop it, once a harness lands. |
| docs / docstrings — the recall tie policy | n/a (a stated policy, not a figure) | PENDING-HARNESS | 'equal relevance => newest first' is measured to FAIL in hybrid and auto modes: RRF gives equivalent records distinct fused scores, so the tie-break never fires and top-1 is the OLDEST. Anywhere the policy is asserted unconditionally it needs scoping to the modes where it holds. |
| bench/ — MemoryAgentBench Conflict Resolution | any CR score | PENDING-HARNESS | Our own gate killed the 'supersession wins CR' headline: a naive keep-all store ties us (97 vs 96, and 85 vs 87 on the faithful re-run), and a 6k-vs-32k context discrepancy in bench/README.md is still open. No CR number may go onto the reader-facing surface until that is resolved. |
| index.html / MCP_LISTINGS.md / README.md — the MCP tool count | 60 | REPRODUCIBLE | No longer hypothetical. Between this audit and its rebase the server grew from 56 to 58 tools; a sibling corrected ONE of the four places that publish the count (the homepage heading) and left the homepage counter, both MCP_LISTINGS figures and the README at 56. This audit named all four on its first run after the rebase. Checked against the live @mcp.tool() count, never typed. |
| bench/README.md — its own committed JSON | 9 of 12 cells | PENDING-HARNESS | Reported to disagree with the JSON it is generated from. Outside this audit's token-enforced scope (README.md, MCP_LISTINGS.md, index.html), and left to the unit that owns the reconciliation rather than guessed at from here. |
| docs / homepage — LongMemEval end-to-end | 0.45, vs a 0.50 oracle ceiling and a 0.05 no-memory floor, n=20 | PENDING-HARNESS | Landed as a pilot. No figure from it has reached README.md, MCP_LISTINGS.md or index.html, and none should until it carries its scope -- n=20, a pilot, and a band check that exits 5 by design. |
What was withdrawn, and why that is the point
Removing a number we cannot back is a win. Every WITHDRAWN row above is a figure this audit
deleted from the reader-facing surface rather than dress up: either no artifact in this repository
produces it, or the artifact it named does not exist. Showing the gap is what makes the rest
worth reading.