Measured results

September 18, 2026 · View on GitHub

Evidence date: 2026-09-18 UTC. Generated offline by bun run bench:report from hash-verified evidence. This is the single results entry point.

Current v12 studies

  • Fable evidence diagnostic (claude-fable-5-1, high effort): Gate disabled; native trigger 1,000 tokens. One trial per arm; strict thinking errors enabled.
  • Sonnet mixed session (claude-sonnet-5, low effort): 50,000-token gate; 190,000 synthetic padding characters; candidate floor 0. One trial per arm.
Study / armInput tokensPeak tokensCache-write tokensCorrectSecondsJev calls / failures
Fable evidence diagnostic / baseline115,45523,78723,755pass39.770 / 0
Fable evidence diagnostic / yoshi75,75013,92412,610pass210.54301 / 3
Fable evidence diagnostic / native108,91410,16776,142pass63.100 / 0
Sonnet mixed session / baseline693,24182,85878,139pass35.140 / 0
Sonnet mixed session / yoshi693,46182,90878,189pass171.99142 / 3
Sonnet mixed session / native664,64075,70783,594pass28.010 / 0
  • Fable evidence diagnostic: 34.39% accumulated input reduction; 41.46% peak reduction. Baseline → Yoshi: 39.77 → 210.54 s; 3 Jev failures. Diagnostic result, not a production-wide savings claim.
  • Sonnet mixed session: -0.03% accumulated input reduction; -0.06% peak reduction. Baseline → Yoshi: 35.14 → 171.99 s; 3 Jev failures. No savings demonstrated.

One trial per arm; models and workloads are separate studies. The Fable gate was disabled. The Sonnet session uses synthetic padding and the default context-size gate, with the candidate floor set to zero. Six Jev failures leave both Yoshi combined costs unknown. No new live Codex benchmark was run. The v11 measurements below are historical.

Input is accumulated provider input, including fresh, cache-read and cache-write tokens. Peak is the largest single request. Cache writes are shown separately because a lower context total does not imply lower cost. Native uses Anthropic tool clearing with two tool uses retained and a 1,000-token minimum clearing amount; its trigger is disclosed per study.

In the strict Fable diagnostic, the proxy replayed real omissions alongside signed thinking without a prefix error or reported thinking drop. The audit checked 92 unchanged-prefix message comparisons. This is transport evidence, not proof that the model used its reasoning. The Sonnet run also passed 92 comparisons, but approved no omissions, so those comparisons do not validate omission replay.

Full validation, development failures and methodology · All 21 development/validation trials

Current cost equivalents

Frozen provider list-price equivalents plus reported Jev cost, not invoices or subscription savings. Missing costs stay unknown. Fable list pricing was not inferred, and missing Jev receipts prevent combined-cost claims for both current Yoshi trials.

StudyBaseline USDYoshi + Jev USDNative USD
Fable evidence diagnosticunknownunknownunknown
Sonnet mixed session0.457019unknown0.472418

Historical v11 studies

Preserved separately; these results do not validate the sticky default. Reference, regression and smoke trials retain their original model, effort and scope.

Study / scenarioModel / effortBaseline inputYoshi inputReductionPeak baseline → YoshiCorrectSeconds baseline → YoshiJev calls / failures
claude-reference / operationsclaude-sonnet-5 / low33,69716,68650.48%33,697 → 16,6862/22.12 → 6.7063 / 0
claude-reference / agent-reportsclaude-sonnet-5 / low54,71721,18961.28%54,717 → 21,1892/23.19 → 9.63107 / 0
claude-reference / evidenceclaude-sonnet-5 / low138,37985,68338.08%25,554 → 11,2772/224.18 → 142.97322 / 3
claude-reference / historyclaude-sonnet-5 / low111,77597,00913.21%17,303 → 15,1642/216.20 → 22.6257 / 0
codex-reference / operationsgpt-6-astra / low33,77422,50333.37%not recorded2/24.73 → 12.3765 / 0
codex-reference / agent-reportsgpt-6-astra / low47,69129,17538.82%not recorded2/25.05 → 49.6684 / 1
codex-reference / evidencegpt-6-astra / low178,642133,28125.39%not recorded2/242.77 → 78.67402 / 0
codex-reference / historygpt-6-astra / low164,805149,3879.36%not recorded2/245.66 → 47.32120 / 0
claude-regression / operationsclaude-sonnet-5 / low33,69028,12316.52%33,690 → 28,1232/23.35 → 46.7318 / 2
codex-smoke / operationsgpt-6-astra / high69,56744,02636.71%not recorded2/28.61 → 21.20146 / 0

Evidence registry

StudyScopeExact answersImmutable evidence
sticky-fable-evidencev12 current3/3Results · Audit
sticky-sonnet-longv12 current3/3Results · Audit
claude-referencev11 reference8/8Results · Audit
codex-referencev11 reference8/8Results · Audit
claude-regressionv11 regression2/2Results · Audit
codex-smokev11 smoke2/2Results · Audit

HTML figures · Machine-readable results · Reproduction · Archived experiments