🦖 stomp report: eval.yaml

September 19, 2026 · View on GitHub

OK: no failures, 1 warning(s) (35 of 35 ran; 67 n/a of 102 declared)

measures the intended construct: NOT ESTABLISHED BY DINOSTOMP

All runs used the offline dry provider; results exercise the benchmark, not any real model.

Results

modelproviderrecordscheckablejudgeableaccuracy95% CIpassesfailsout tokspend
dry-alphadry2424100%100.0%[0.862, 1.000]24024$0.0000
dry-bravodry2424100%75.0%[0.551, 0.880]18624$0.0000
dry-charliedry2424100%37.5%[0.212, 0.573]91524$0.0000
dry-deltadry2424100%100.0%[0.862, 1.000]24024$0.0000
dry-echodry2424100%50.0%[0.314, 0.686]121224$0.0000
dry-foxtrotdry2424100%54.2%[0.351, 0.721]131124$0.0000

Accuracy is ON CHECKABLE output: judgeable is the share the scorer reached a verdict on at all, and 80% accurate on 60%-judgeable output is not 80% accurate.

6 model(s) x 24 item(s), mean 69.4%, spanning 37.5% to 100.0% (62% spread), KR-20 0.94.

9 item(s) every model passed and 0 every model failed: 38% of the set separated nobody in this fleet.

At 24 items an UNPAIRED comparison resolves gaps down to about 40%; smaller differences between the models above are not distinguishable from sampling noise by that test.

Item difficulty: all 24 item(s), hardest first
itemtargetpdiscriminationmissed bymost common wrong answer
a132733%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot28
a224533%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot47
a255133%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot54
a306133%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot66
a316333%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot66
a336733%+0.87dry-bravo, dry-charlie, dry-echo, dry-foxtrot71
a142950%+0.90dry-charlie, dry-echo, dry-foxtrot36
a183750%+0.90dry-charlie, dry-echo, dry-foxtrot42
a193950%+0.90dry-charlie, dry-echo, dry-foxtrot44
a204150%+0.90dry-charlie, dry-echo, dry-foxtrot43
a295950%+0.90dry-charlie, dry-echo, dry-foxtrot61
a265367%+0.71dry-charlie, dry-echo58
a102183%+0.54dry-charlie28
a122583%+0.54dry-charlie29
a153183%+0.54dry-charlie33
a1123100%---
a1633100%---
a1735100%---
a2143100%---
a2347100%---
a2449100%---
a2755100%---
a2857100%---
a3265100%---

p is the share of the fleet that answered correctly and discrimination is the point-biserial with fleet skill. Both DESCRIBE; a hard item is not a defect. A negative discrimination is what P2 examines.

Entitled claims

This result is entitled to claim:

  • Exact-match accuracy with a 95% interval on these 24 addition items, bare-number format, per model.

Typed claims, compiled to evidence requirements and checked off:

  • SUPPORTED: accuracy of dry-alpha is at least 80% (95% confidence)
    • complete run on disk: dry-alpha: complete
    • enough checkable evidence: 24 checkable unit(s); need 20
    • interval lower bound clears the declared minimum: lower bound 86.2% vs declared minimum 80%
  • SUPPORTED: dry-alpha beats dry-charlie by at least 20% (95% confidence)
    • complete runs for both models: dry-alpha: ok; dry-charlie: ok
    • paired observations: 24 common item(s); need 20
    • paired bootstrap clears min_effect at the declared confidence: gap >= 20% in 100% of 400 resamples; need 95%

Checks

Invariants (deterministic, gating)

Facts, not heuristics: a failure here means something is mechanically wrong (a duplicate exists, a hash changed, a number does not re-derive) and it breaks the verdict.

checkwitnessesdetail
okquestions are unique240 duplicated question(s) among 24
okno answer leaks into its own question240 of 24 item(s) leak their answer into the question
n/ano option offered twice in one item0no multiple-choice items in this dataset
n/aevery target is among its choices0no multiple-choice items in this dataset
okno identical question with contradictory targets240 question(s) appear with conflicting targets
n/aevery referenced asset resolves and still hashes the same0no item carries an input_ref; nothing points at a file
n/ano asset's own path gives away its label0no item carries an input_ref; nothing points at a file
n/ano asset appears in two splits0no item carries an input_ref; nothing points at a file
okthe audit covers the rows it was given240 of 24 row(s) were dropped: the pod loader refuses a dataset it cannot read whole, so every row in the file reached the audit
n/arows are unique0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano error value is saved in the workbook0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/aevery column aggregate covers its own column0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/athe join returns rows at all0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano key fails to match on case or whitespace alone0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/aevery parent total equals the sum of its children0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/aa graded scorer witnesses its gradation0this scorer does not emit intermediate partial credit, so there is no gradation to witness
okevery typed claim's evidence requirements hold62 of 2 typed claim(s) supported across 6 evidence requirement(s) (no multiplicity correction across 2 claims)
okruns match the spec, data, and scorer on disk (no drift)60 of 6 run(s) no longer match the spec, data, or scorer on disk
okthe witness gate replays clean11replayed 5 witness(es): 5 behaved; 6 run manifest(s) checked
okledger spend agrees with the manifest and the spec cap60 money discrepanc(ies) across 6 run(s)
okevery run record is schema-valid, unique, and its manifest's own1440 integrity problem(s) across 144 record(s)
oktruncated outputs are never credited1440 truncated output(s) scored as pass; a cut-off response can still have stated its answer, so read these before raising max_tokens and re-running
okrecorded verdicts re-score identically1440 of 144 recorded verdict(s) do not reproduce under the current scorer
oksummaries match their run records60 summary discrepanc(ies) across 6 run(s)
okrecords cover exactly the seeded selection60 of 6 run(s) do not cover their seeded selection
okevery model produced something scoreable60 of 6 model(s) produced nothing scoreable
n/agraded scores stay in range0no record carries a graded value
n/ano forbidden tool is called0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/aevery required tool is actually called0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/atrajectories are well-formed0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
okevery model was asked the same items60 of 6 model(s) were asked a different item set

Diagnostics (statistical, advisory)

Threshold-based signals: they warn, expose their underlying values, and can have legitimate explanations. A warning is evidence of possible trouble, never a proof of invalidity.

checkwitnessesdetail
n/agold answer does not favour an option position0no multiple-choice items in this dataset
n/agold answer is not systematically the longest option0no multiple-choice items in this dataset
oka contamination canary travels with the data1canary present (dinostomp canary DO NOT TRAIN 4e3844fb6f...)
n/ano surface feature predicts the gold answer0no multiple-choice items in this dataset
n/ano model reproduces the contamination canary0regurgitation probes need a hosted model; this pod's runs are all local
n/ano item already appears in a reference dataset0no reference dataset supplied; pass --against to compare these items against a corpus you have. This never checks training data, and cannot.
n/ano near-duplicate assets0no item carries an input_ref; nothing points at a file
n/athe eval is not authored in a circle0no provenance declared, so authorship is not described. Declaring who wrote the items, keys, scorer, and witnesses lets this surface a model sitting on both sides of a loop (e.g. keying its own questions)
n/ano single column all but determines the target0an eval pod's items are questions and answers, not a feature table; the single-column leak scan is for a raw tabular dataset audit
n/ano two options are the same number written differently0no multiple-choice items in this dataset
okno two items are the same question in different encodings240 group(s) of items are the same question in different encodings
okthe answer key is not dominated by one value24the answer key is not dominated by one value (modal 4% of 24 answers)
n/ano cell carries edge whitespace or invisible characters0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano identifier column repeats a value0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano digit-string column has leading zeros a conversion would destroy0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano date column mixes formats or reads both ways0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano numeric column is contaminated with text0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano value stands in for missing without saying so0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano category column splits one label across spellings0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano rate column mixes fraction and percent scales0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano numeric column stores symbols or separators0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano rows are duplicates once case and whitespace stop counting0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano constant is pasted inside a formula column0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/anothing an aggregate counts is hidden from the reader0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano merged range flattens a row on import0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/aevery formula has been calculated0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/ano left row is dropped by the join0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/athe right-hand key is unique0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/athe join does not multiply rows0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
n/aboth sides store the key the same way0out of scope for a pod audit (the dataset, the scorer, and every run on disk)
okwitnesses kill the mutant scorers50 of 5 applicable mutant scorer(s) survive the witness suite
n/aa correct answer survives its surface form0this scorer compares exactly rather than extracting, so surface-form robustness is not a property it claims
okan exact scorer is not graded against prose answers24answers are short (1-word median); exact match fits
okuncheckable rate is sane1440% of 144 record(s) are uncheckable
okaccuracy is distinguishable from guessing60 of 6 model(s) score no better than guessing; fleet spans 38% to 100% vs chance ~4% (modal target floor)
okruns cover the spec's declared scope, nothing foreign60 run(s) outside the spec's declared scope
okno model selectively escapes the scorer60 of 6 model(s) escape the scorer more than the fleet does
n/athe eval is not solvable blind0blind probes need a real provider; this pod's runs are all dry
okno model collapses onto one answer60 of 6 model(s) answer with one response far more often than any target warrants
n/aeach model beats its own blind baseline0blind probes need a real provider; this pod's runs are all dry
okfailed answers do not contain the reference40 of 4 model(s) are failed on answers that contain the reference; the scorer may be grading format, not correctness
n/abilled output tokens match the recorded text0no model produced 20+ answers of at least 40 characters; short-answer evals cannot be billed against reliably
warnthe runs were produced by this engine66 of 6 run(s) were produced by a different engine than the one auditing them (now 203b45ecbf3cceee); re-run to get numbers this report can stand behind
n/arepeated items reached a verdict0no run on disk repeats an item; a single pass per item cannot tie
okno failed answer numerically equals its target440 of 44 numeric-target failure(s) equal their target as a number; the scorer may be rejecting a correct value in the wrong form
n/areported confidence matches observed accuracy0no record carries a probability vector; only a one-pass model (decisions, chooser, loglikelihood) reports one
n/aconfidence separates right answers from wrong ones0no record carries a probability vector; only a one-pass model (decisions, chooser, loglikelihood) reports one
n/apassing answers are grounded in tool evidence0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/ano model under-reports its trajectory0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/atool calls are not redundant0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/apassing answers CHANGE when their evidence is withheld0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/athe trajectory was observed, not self-reported0this spec runs no code targets and no imported run carries a trajectory; nothing here produces or carries one
n/aa timed-out call that already ran is not run again0only a mediated agent reaches its tools through the harness, so there is no call to time out
n/athe judge agrees with cases whose answer is known0this eval does not score with a judge
n/athe judge is invariant to content-free perturbations0this eval does not score with a judge
n/athe judge agrees with itself on identical input0this eval does not score with a judge
n/athe judge does not favour its own family0this eval does not score with a judge
okfleet score totals are reliable (KR-20)144KR-20 0.94 across 6 models x 24 items; small fleet (6 examinees), treat as a noisy estimate
okno item anti-correlates with fleet skill240 item(s) that strong models miss and weak models hit, against 0 expected by chance at this fleet size; candidate key errors; at 6 examinees this check has little power, so a quiet result is NOT evidence of a clean answer key
okdead-weight items stay a minority2438% of 24 item(s) separate nobody (9 all-right, 0 all-wrong); 8% would be dead at 6 examinees even with no difficulty structure, so part of this is fleet size
okno unanimous identical wrong answers240 item(s) where the whole fleet gave one identical wrong answer; candidate key errors
n/aentitled ordering claims are separated beyond sampling noise0no entitled claim asserts a model ordering
okthe fleet is not pinned at a ceiling or floor6fleet accuracy spans 38% to 100% on 24 item(s)
okthe eval separates the fleet (dynamic range)6fleet spread 62% across 6 model(s) on 24 item(s)
n/aanswers survive re-ordering the options0presentation-order probes need a real provider; this pod's runs are all local
n/athe number survives changing the seed0the spec declares no extra seeds; a single seed cannot show its own spread (run.seeds is how you ask)
n/athe number survives re-phrasing the instruction0instruction-framing probes need runs on disk
n/athe fleet ORDERING survives re-phrasing the instruction0instruction-framing probes need runs on disk
okthe fleet varies on one axis, not a blend of abilities24top-axis share 0.74 against 0.84 the fixed-margins null allows; the fleet varies on one axis; this fleet is all-dry, whose skill is a single scalar by construction, so a quiet result here is a plumbing check, not validity evidence; at 6 examinees this has limited power, so a quiet result is NOT proof the score measures one thing
n/adeclared subskills actually separate in the responses0no item declares a subskill; there is no partition to test
n/aanswers survive a changed tool list0menu probes need a real provider; this pod's runs are all local

Re-derive this report from the directory holding the target: dinostomp stomp eval.yaml

Receipts

[ok] the answer key is not dominated by one value
  • evidence: {"modal_share": 0.042, "modal_value": "21", "n_distinct": 24}
[ok] the audit covers the rows it was given
  • evidence: {"dropped_share": 0.0, "gate": 0.01, "rows_audited": 24, "rows_dropped": 0, "rows_read": 24}
[ok] witnesses kill the mutant scorers
  • evidence: {"killed": ["always-pass", "always-fail", "substring-lenient", "prefix-lenient", "negation-blind"], "not_applicable": ["case-blind", "space-blind", "uncheckable-credit"]}
[ok] an exact scorer is not graded against prose answers
  • evidence: {"long_share": 0.0, "median_answer_words": 1}
[ok] uncheckable rate is sane
  • evidence: {"rate": 0.0}
[ok] accuracy is distinguishable from guessing
  • evidence: {"chance_floor": 0.0417, "modal": 0.0417, "modal_target": "21", "per_model_accuracy": {"dry-alpha": 1.0, "dry-bravo": 0.75, "dry-charlie": 0.375, "dry-delta": 1.0, "dry-echo": 0.5, "dry-foxtrot": 0.5417}, "uniform": 0.0}
[ok] no model selectively escapes the scorer
  • evidence: {"rates": {"dry-alpha": 0.0, "dry-bravo": 0.0, "dry-charlie": 0.0, "dry-delta": 0.0, "dry-echo": 0.0, "dry-foxtrot": 0.0}}
[warn] the runs were produced by this engine
  • engine 050e2f343915e1b9: dry-alpha seed 42 (tool 0.57.1), dry-bravo seed 42 (tool 0.57.1), dry-charlie seed 42 (tool 0.57.1) and 3 more
  • evidence: {"engines": {"050e2f343915e1b9": 6}}
[ok] fleet score totals are reliable (KR-20)
  • evidence: {"excluded_collapsed": [], "kr20": 0.9443, "n_examinees": 6}
[ok] no item anti-correlates with fleet skill
  • evidence: {"chance_95th": 0, "excluded_collapsed": [], "n_examinees": 6, "negative_rpb": 0, "underpowered": true}
[ok] dead-weight items stay a minority
  • evidence: {"independence_floor": 0.0762, "n_examinees": 6, "share": 0.375}
[ok] the fleet is not pinned at a ceiling or floor
  • evidence: {"max": 1.0, "min": 0.375}
[ok] the eval separates the fleet (dynamic range)
  • evidence: {"spread": 0.625}
[ok] the fleet varies on one axis, not a blend of abilities
  • evidence: {"all_dry": true, "concentration": 0.7438, "margin": 0.1, "n_examinees": 6, "null_95": 0.7438}

Runs

run filemodelreported asproviderdryseedrecordsuncheckable
20260810_110453_fleet-arith_dry-alpha_n24_s42.jsonldry-alpha(same)dryyes42240
20260810_110453_fleet-arith_dry-bravo_n24_s42.jsonldry-bravo(same)dryyes42240
20260810_110453_fleet-arith_dry-charlie_n24_s42.jsonldry-charlie(same)dryyes42240
20260810_110453_fleet-arith_dry-delta_n24_s42.jsonldry-delta(same)dryyes42240
20260810_110454_fleet-arith_dry-echo_n24_s42.jsonldry-echo(same)dryyes42240
20260810_110454_fleet-arith_dry-foxtrot_n24_s42.jsonldry-foxtrot(same)dryyes42240

Provenance

  • tool: dinostomp 0.63.0
  • statistical power: at n=24 items, an UNPAIRED comparison (worst case p=0.5) resolves gaps down to ~40% accuracy (80% power, two-sided alpha 0.05); the paired bootstrap behind P6/C1 resolves smaller gaps when model errors overlap
  • spec_sha256: cc280622dc0f91aa3809e0072bf4125add1a99825319f4f97a681fb4e23657cc
  • data_sha256: 742d30fda48436edace090596fe7659588d02761e788be1acdd3fbf2573437dd
  • thresholds: all defaults
  • reproducibility tiers, stated honestly: local inputs hash-pinned (spec, data, scorer); requests reproducible given each manifest's environment envelope; hosted-model immutability UNKNOWN unless the provider exposes a pinned revision (the runs table records what each provider claims answered)
  • raw report: STOMP.json (both files omit volatile fields, so an unchanged pod re-reports to identical bytes; run manifests carry the timestamps)