Release implementation verification, frozen before execution
September 20, 2026 ยท View on GitHub
Compare the compiled 0.2.0 jev_label MCP server against plain Codex on the same 64-message synthetic feedback dataset used during development. This is an implementation-transfer check on known development data, not a held-out or statistical confirmation. The original positive prototype result and all preceding negative phases remain in the historical evidence.
Four paired repetitions (eight sessions), alternating A,N and N,A. Both use gpt-6-astra, medium reasoning, priority service, identical input, policy and full artifact requirements, ordinary shell tools, and fresh temporary task folders. The plugin uses eight concurrent batches of up to eight Jev Choice questions, pinned jev-1.13.0, and confidence threshold 0.8. A benchmark-only fetch observer records status, timing, usage and raw labels without headers or credentials. It is not part of the production launcher. Only the writer tool is approved in the isolated benchmark; global permissions are unchanged.
Before the first session, freeze source, compiled files, runner, observer, grader, protocol, inputs and gold. Grade the whole output artifact for exact coverage, fields and labels. Record actual provider usage, raw errors, high-confidence errors, final errors, commands and final time/token usage. Keep every run, including failures and incidental command errors. Do not resume or overwrite an existing output directory, selectively retry a slow/incorrect run or stop early when results look favorable.
The unchanged gate requires all intervention outputs correct, actual Jev use, at least 20% less paired geometric mean total Codex input or elapsed time, with the other measure no more than 10% worse. Report all-assigned results first. A secondary post-hoc analysis may exclude entire pairs affected by a failed shell command, clearly labelled as sensitivity. Never remove only an inconvenient arm.
Synthetic labels are authored by the same agent and the clean balanced messages often state the primary request explicitly. Repetitions do not increase the 64 unique-record sample. The confidence threshold is not calibrated, and high confidence can hide mistakes. Total input includes cached tokens; separate uncached/output and provider usage are required. No invoice, subscription, precision superiority, general labelling or desktop auto-discovery claim follows from a passing development gate. Representative independently labelled data is needed before those claims.