Measured results
September 18, 2026 · View on GitHub
Evidence date: 2026-09-18 UTC. Generated offline by bun run bench:report from hash-verified evidence. This is the single results entry point.
Current v12 studies
- Fable evidence diagnostic (claude-fable-5-1, high effort): Gate disabled; native trigger 1,000 tokens. One trial per arm; strict thinking errors enabled.
- Sonnet mixed session (claude-sonnet-5, low effort): 50,000-token gate; 190,000 synthetic padding characters; candidate floor 0. One trial per arm.
| Study / arm | Input tokens | Peak tokens | Cache-write tokens | Correct | Seconds | Jev calls / failures |
|---|---|---|---|---|---|---|
| Fable evidence diagnostic / baseline | 115,455 | 23,787 | 23,755 | pass | 39.77 | 0 / 0 |
| Fable evidence diagnostic / yoshi | 75,750 | 13,924 | 12,610 | pass | 210.54 | 301 / 3 |
| Fable evidence diagnostic / native | 108,914 | 10,167 | 76,142 | pass | 63.10 | 0 / 0 |
| Sonnet mixed session / baseline | 693,241 | 82,858 | 78,139 | pass | 35.14 | 0 / 0 |
| Sonnet mixed session / yoshi | 693,461 | 82,908 | 78,189 | pass | 171.99 | 142 / 3 |
| Sonnet mixed session / native | 664,640 | 75,707 | 83,594 | pass | 28.01 | 0 / 0 |
- Fable evidence diagnostic: 34.39% accumulated input reduction; 41.46% peak reduction. Baseline → Yoshi: 39.77 → 210.54 s; 3 Jev failures. Diagnostic result, not a production-wide savings claim.
- Sonnet mixed session: -0.03% accumulated input reduction; -0.06% peak reduction. Baseline → Yoshi: 35.14 → 171.99 s; 3 Jev failures. No savings demonstrated.
One trial per arm; models and workloads are separate studies. The Fable gate was disabled. The Sonnet session uses synthetic padding and the default context-size gate, with the candidate floor set to zero. Six Jev failures leave both Yoshi combined costs unknown. No new live Codex benchmark was run. The v11 measurements below are historical.
Input is accumulated provider input, including fresh, cache-read and cache-write tokens. Peak is the largest single request. Cache writes are shown separately because a lower context total does not imply lower cost. Native uses Anthropic tool clearing with two tool uses retained and a 1,000-token minimum clearing amount; its trigger is disclosed per study.
In the strict Fable diagnostic, the proxy replayed real omissions alongside signed thinking without a prefix error or reported thinking drop. The audit checked 92 unchanged-prefix message comparisons. This is transport evidence, not proof that the model used its reasoning. The Sonnet run also passed 92 comparisons, but approved no omissions, so those comparisons do not validate omission replay.
Full validation, development failures and methodology · All 21 development/validation trials
Current cost equivalents
Frozen provider list-price equivalents plus reported Jev cost, not invoices or subscription savings. Missing costs stay unknown. Fable list pricing was not inferred, and missing Jev receipts prevent combined-cost claims for both current Yoshi trials.
| Study | Baseline USD | Yoshi + Jev USD | Native USD |
|---|---|---|---|
| Fable evidence diagnostic | unknown | unknown | unknown |
| Sonnet mixed session | 0.457019 | unknown | 0.472418 |
Historical v11 studies
Preserved separately; these results do not validate the sticky default. Reference, regression and smoke trials retain their original model, effort and scope.
| Study / scenario | Model / effort | Baseline input | Yoshi input | Reduction | Peak baseline → Yoshi | Correct | Seconds baseline → Yoshi | Jev calls / failures |
|---|---|---|---|---|---|---|---|---|
| claude-reference / operations | claude-sonnet-5 / low | 33,697 | 16,686 | 50.48% | 33,697 → 16,686 | 2/2 | 2.12 → 6.70 | 63 / 0 |
| claude-reference / agent-reports | claude-sonnet-5 / low | 54,717 | 21,189 | 61.28% | 54,717 → 21,189 | 2/2 | 3.19 → 9.63 | 107 / 0 |
| claude-reference / evidence | claude-sonnet-5 / low | 138,379 | 85,683 | 38.08% | 25,554 → 11,277 | 2/2 | 24.18 → 142.97 | 322 / 3 |
| claude-reference / history | claude-sonnet-5 / low | 111,775 | 97,009 | 13.21% | 17,303 → 15,164 | 2/2 | 16.20 → 22.62 | 57 / 0 |
| codex-reference / operations | gpt-6-astra / low | 33,774 | 22,503 | 33.37% | not recorded | 2/2 | 4.73 → 12.37 | 65 / 0 |
| codex-reference / agent-reports | gpt-6-astra / low | 47,691 | 29,175 | 38.82% | not recorded | 2/2 | 5.05 → 49.66 | 84 / 1 |
| codex-reference / evidence | gpt-6-astra / low | 178,642 | 133,281 | 25.39% | not recorded | 2/2 | 42.77 → 78.67 | 402 / 0 |
| codex-reference / history | gpt-6-astra / low | 164,805 | 149,387 | 9.36% | not recorded | 2/2 | 45.66 → 47.32 | 120 / 0 |
| claude-regression / operations | claude-sonnet-5 / low | 33,690 | 28,123 | 16.52% | 33,690 → 28,123 | 2/2 | 3.35 → 46.73 | 18 / 2 |
| codex-smoke / operations | gpt-6-astra / high | 69,567 | 44,026 | 36.71% | not recorded | 2/2 | 8.61 → 21.20 | 146 / 0 |
Evidence registry
| Study | Scope | Exact answers | Immutable evidence |
|---|---|---|---|
| sticky-fable-evidence | v12 current | 3/3 | Results · Audit |
| sticky-sonnet-long | v12 current | 3/3 | Results · Audit |
| claude-reference | v11 reference | 8/8 | Results · Audit |
| codex-reference | v11 reference | 8/8 | Results · Audit |
| claude-regression | v11 regression | 2/2 | Results · Audit |
| codex-smoke | v11 smoke | 2/2 | Results · Audit |
HTML figures · Machine-readable results · Reproduction · Archived experiments