Who&When Pro - Jev vs gpt-5.4 (text subset)

September 17, 2026 ยท View on GitHub

ModelWhoWhenWhatAll
Jev (ours, n=6257)73.4 [70.5, 76.2]76.4 [74.3, 78.4]23.7 [22.4, 24.9]31.3 [29.5, 33.0]
gpt-5.4 (paper)55.772.315.321.3

Jev cost: $1.28 for all 6257 traces (30,497,481 input tokens at $0.042/Mtok; output free).

Who and When are adaptation-favoured: Jev picks the agent/step from the trace's enumerated options, while the LLM free-generates. The like-for-like axis is What (error macro-F1, same 17-code taxonomy for both). Mode-confidence ECE 0.287. Latency is not compared (shared early-access endpoint). Baseline: gpt-5.4, arXiv:2607.09996 Table 4. Failures in Who&When Pro are injected, not natural incidents.