Research log

September 20, 2026 · View on GitHub

Results are tied to implementation hashes, populations, and declared baselines. Keep negative results and coverage gaps. A later change does not inherit an earlier version's qualification automatically. Git history preserves prior reports; each entry below names the work and its limits.

2026-09-20 UTC — catalog scorecard now separates results from demand screens

The public catalog now gives every entry a machine-readable decision and keeps the metrics on separate axes. system-one-verify remains the only shipped skill with a measured result: 35.20% less presented validation text across 563 replayed outputs, with zero preservation failures. It is a text-boundary replay, not provider billing or complete-task savings.

The other workflow numbers are explicitly headroom screens, not reductions: 213/573 repository-search calls (37.17%), 31/68 diff-review calls (45.59%), 28/224 Git-state calls (12.50%), and 54/95 shared web-research calls (56.84%) were at least 256 o200k_base output tokens. That threshold assumes a hypothetical 128-token skill cost and 128-token margin, then imagines perfect deletion. It does not establish that an adapter exists, preserves the answer, or beats the native tool. Fetch and research share the web denominator and are not two results. The CI pilot remains native-baseline preferred; triage, writing, evolve, and the umbrella router have no dedicated labeled cohort. Scorecard and derivation.

2026-09-20 UTC — exact read reuse did not justify another skill

A new local screen covered all 1,200 file reads in the fixed discovery window: 42 Codex, 9 Claude, and 1,149 Devin calls. It found 74 exact repeated responses, containing 12,485 of 1,143,648 read-output text tokens (1.09%). This is an optimistic deletion ceiling before overhead, not measured savings. The stricter subset with no intervening recorded tool call was one Claude response worth two tokens. No read was established to be safe to skip, and read reuse remains outside the package. Results, provider counts, and limits.

Explicit native read bounds already appeared in 516 calls (43%). Zero complete multi-entry JSON output fragments met the narrow field-projection format screen. Unrecognized bounds and absent JSON matches are not evidence of waste elsewhere. The collection inspected retained local transcripts only; it made no provider requests and executed no historical commands. The local protocol preceded the new outcomes, and private opaque measurements bind the public rows to exact query, scope, output, and chronology hashes. Post-collection independent review hardened only report validation; the original source hash and amendment are retained.

The installed skill and runtime are unchanged. The result narrows the research queue instead of adding another overhead-bearing instruction without evidence.

2026-09-20 UTC — Codex comparison completed after setup repairs

Two fresh Codex CLI runs diagnosed the same real Devin failure log. Both answers passed all six criteria in an AI review blinded to arm labels and token usage. The compact-result arm used 47,840 recorded input-plus-output tokens; focused native reads used 49,803. The descriptive difference is 1,963 (3.94%). The compact arm made four read commands versus three and took 37,974 ms versus 28,625 ms from launch to process exit. Neither better reliability nor faster completion is demonstrated. Plain-language results.

The original startup timeout remains unchanged. A first retry using initialized default state started successfully but exposed a disabled tool-host setting; its 24,293 recorded tokens and unusable result remain in a separate report. Enabling the installed stable tool host allowed the completed pair. All three retry launches total 121,936 recorded tokens; original-timeout usage remains unknown. No failed attempt is removed or assigned an invented zero cost.

Both setups had private plans frozen before their launches and later public protocol projections. Inputs, schema, skill, model request, and rubric remained fixed. New reports bind the plan, captures, and locked blind review. One case, fixed order, cache differences, unversioned model identity, and imperfect isolation attestation prevent causal or general benefit claims. No new skill or expanded use case is admitted; the v0.4.0 runtime remains unchanged.

2026-09-20 UTC — approved diagnosis attempt timed out

The owner explicitly authorized sending the selected private Devin failure log to Codex for the two-session diagnosis comparison. The first launch, in the frozen reduced-then-native order, timed out after 120 seconds with no output, answer, or usage counters. Read-only inspection found startup history import still running and no existing managed daemon for the supported proxy route. The experiment stopped before the native arm under its startup stop rule.

The observed-attempt report preserves one launch, zero completed answers, unknown upstream request status, and missing token measurements. Launcher duration is not diagnosis latency. The original pre-approval readiness record remains unchanged; it describes the earlier rejected attempt only. See the pilot chronology and next requirements.

A new deterministic assessor checks launch order and inventory, source hashes, token counter accounting, and correctness-review provenance. Its synthetic tests verify the assessment logic; they supply no new skill-efficacy observations. No savings, reliability, or speed claim changes, and the v0.4.0 runtime remains unchanged. A usable startup environment and a separately recorded continuation plan are needed before another paired experiment can produce evidence.

2026-09-20 UTC — failure evidence and native alternatives

One real noisy failure: a preselected retrospective failure scan found one qualifying Devin excerpt and none for Codex or Claude. The frozen reducer saved 4,253 text tokens (66%) after counted first-use overhead. It retained all seven failed-test identities and their shared error message, but omitted six of seven distinct test callsites. The labels were independently checked against the raw excerpt before replay. One complete later log read would erase the saving; that is a cost sensitivity, not an observed agent action. A separate color-stripped text replay saved 2,118 tokens (58%). No diagnostic-success or speed claim follows. Failure audit.

Four candidate workflows: exact text counts for 960 retained call outputs from Codex, Claude, and Devin show that many results are already small. Even deleting all output cannot meet a hypothetical 64-token cost plus 128-token savings target for 57% of search calls, 81% of Git-state calls, 51% of diff calls, and 41% of web calls. These are optimistic output-only ceilings, not measured candidate savings, and reuse the discovery period rather than a fresh holdout. They do not evaluate fewer reasoning turns or future calls. Candidate screen.

An actual native alternative: this repository's existing 21-test assessment file returned the same passing counts with normal, dots, and failures-only Bun reporters. Output plus invocation text fell from 466 tokens to 55 or 59 without a skill. The short passing check uses synthetic test inputs and supplies no historical-project, warning-preservation, diagnosis, or latency result. The three older favorable calibration logs contain Bun test output; native quiet reporters were not compared on their original snapshots. That limitation is now explicit beside the headline's calculation.

Earlier pre-approval attempt: a one-case paired Codex diagnosis pilot was prepared, but no model session started in that attempt. Runtime approval review rejected transmission of the private excerpt; no payload was sent and no alternate route was attempted. The pilot status records zero completed pairs and null observed metrics. The later authorized attempt is recorded above. Local text analysis cannot substitute for that trial.

Distribution unchanged: one shipped skill and ten uninstalled candidates. The catalog now leads with the decision for each skill rather than repeated empty metric columns. The v0.4.0 runtime and installed instructions are unchanged; public evidence checks cover the new source hashes, selection counts, arithmetic, and unfavorable outcomes.

2026-09-19 — clearer impact summary

The README and plain-language results now lead with 82% fewer tokens for noisy check results, scoped immediately to an initial replay of three successful Devin logs. This is the existing aggregate result expressed as a percentage: (9,731 − 931 − 774) / 9,731 = 82.48%. It includes counted skill overhead. No new trial, broader efficacy claim, or runtime change is implied. Technical methods and unfavorable cases remain in the linked reports. The next evidence needed is complete unused noisy-check tasks against strong native baselines, not a larger raw transcript count alone.

2026-09-19 — full inventory and separate evidence axes

Distribution: one shipped skill, system-one-verify; ten research candidates. The complete catalog lists each proposed mechanism, observed transcript coverage, stronger native alternative, correctness contract, missing costs, and next experiment. Listing a candidate does not install it or claim that it works better.

Token evidence: the v0.4 calibration report retains all 24 selected Devin excerpts. Three qualifying logs save 8,026 o200k_base text tokens after 774 static instruction/invocation tokens. Twenty-one short cases would add cost. No whole-task or billed-token conclusion follows. The two calibration cohorts also contain real Codex and Claude records, but no eligible validation replay samples from those providers. Methods and per-case results.

Performance evidence: 140 balanced native/wrapper pairs over seven transcript-shaped synthetic fixtures measured 41.82 ms median / 48.37 ms p95 added local wall time. All timings, warmup policy, source hashes, and shared macOS/Node environment are published. This is overhead, not a task speedup. Runtime results.

Correctness evidence: the process measurement verified exact exits, one execution, full private logs, and byte-exact short output. Twenty-five runtime tests cover the documented fault and capture contracts. No trial measured a statistical improvement in agent task reliability or diagnostic quality. The distinction is explicit in the README and catalog.

Assessment improvement: the whole-task evaluator now requires audited native baselines, matched context, reconciled token buckets, all attempts and follow-ups, blinded outcome evaluation, and unused tasks. Shared task/session ancestry reduces the independent sample count. Correctness regressions block every positive adoption screen; material whole-task latency regressions cannot be offset by token savings. The program reports separate verdicts and uncertainty. There are no qualifying whole-task observations yet; unit fixtures are not presented as real trials.

New historical cohort: a separate September 1–11 window tests the frozen runtime and threshold against previously unused archived text. The locally recorded protocol precedes collection; it is not an independently timestamped public preregistration. The first collection failed on non-object Claude tool metadata before a complete corpus or replay was produced. A recorded amendment adds a separately hashed adapter that normalizes unsupported metadata to an empty mapping and counts it. It neither infers exit statuses nor changes the frozen original analyzer or calibration reports. Raw logs and replay artifacts stay private; archived commands never run.

The amended collection contains 35,307 tool calls (86 Codex, 33,710 Claude, 1,511 Devin). Its deterministic selection yields 14 completed validation replays: two Claude and twelve Devin. All are below 8 KiB, so none qualifies for the wrapper. Invoking it on every case would add 3,612 estimated tokens; ten cases have recorded nonzero exits and no preservation invariant fails. This confirms no new noisy-log savings and supplies no basis to expand the shipped skill. Complete results, exclusions, and reproduction. Integrity checks accept this unfavorable outcome rather than require a positive result. The expanded provider corpus must not be confused with eligible replay coverage or a representative task sample.

CI candidate pilot: an identity audit of the 128 Devin CI-status calls resolves 32 calls into 30 explicit run/context groups; 96 remain unresolved. Only two groups repeat, with different query options in each. Neither group establishes avoidable model polling, token savings, or an improvement over native run watching. The candidate remains uninstalled. Results and conservative parser exclusions.

Next admission work

  1. Test validation on independent paired complete tasks, including long failures, wrong skill selection, warnings/coverage questions, and native quiet reporters. Use actual provider usage and count the cost of every follow-up log read.
  2. Investigate exact-location repository search: the calibration corpus has 573 search calls across all three providers. Compare against focused native search with known correct answer locations. Frequency alone proves no benefit.
  3. Before further CI implementation, obtain task-intent evidence for genuinely repeated unchanged-status checks and compare with native run watching. The current identity pilot does not establish such a case; frequency alone is insufficient reason to add a wrapper.

Each experiment must freeze scope and criteria before collecting the data used for its claim. Publish an insufficient or negative result as readily as a positive one. Promote only the supported workload, withdraw stale claims, and keep the default skill footprint small.