Where this suite is weak

September 2, 2026 ยท View on GitHub

A conformance suite that publishes only what it covers is asking to be taken on trust. This document is the other half: the conditions this corpus cites and does not uniquely force, the failure codes a conforming verifier can decline to emit entirely, and the sites no vector can ever reach. Every number here is derived from the published mutation baseline on every run, by the gate named in the comment above, so it cannot drift from the data it summarises.

The figure, and what it measures

Of the 93 normative conditions this corpus cites, 66 are forced by a vector no other condition's vectors duplicate, 27 are covered only redundantly, and 0 are not forced at all.

whatvalue
suiteadversarial-execution-evidence-conformance
corpussuiteRevision 28, 272 vectors (61 accept, 209 reject, 2 indeterminate)
vendored specificationcommit 0dbe10bcc959b63dc42370a5db09812c9476f59a, fetchable from astrogilda/attestation at branch predicate/adversarial-execution-evidence
reviewed atin-toto/attestation#570
vectors/MANIFEST.jsonsha256:7b6c7cc5f0068f5070f0da1eef2d43387eb070858cfd4041f850674de29471f0
docs/FORCING-BASELINE.jsonsha256:8a108f068cb98418661a9a84b4b94dbf9655a99d8cb4be7c4f537cae812de5fb
campaign807 single-site weakenings: 459 KILLED, 32 SILENT, 311 DEAD, 5 INCONCLUSIVE

Every vector in the corpus is depended on by at least one recorded weakening, so no vector is wholly redundant and the campaign covers the corpus as it stands.

The corpus cites the specification's rules by aee-c-NN condition id on every vector. Take the vectors citing one condition, and ask what the mutation campaign in docs/FORCING-BASELINE.json says about them:

  • forced -- some rail weakening is caught by those vectors and by no others. Delete them and that weakening goes undetected.
  • weak -- they catch weakenings, but every one is also caught by a vector citing something else. The rule is covered; nothing is uniquely attributable to this condition.
  • not forced -- they catch no weakening at all.

Weak is a statement about attribution, not about strength, and it does not move in one direction. Adding a vector that happens to catch an already-caught weakening enlarges that weakening's killer set, which can move a condition from forced to weak while the corpus got strictly better. Two of these figures from different revisions are not a trend line, and a rise in the weak count is not a regression.

The figure this replaces, and what could not be re-derived about it. A condition projection was quoted before this page existed: 24 of 78 conditions weak, over the corpus at suiteRevision 14. Its own campaign put its site-axis half in vectors/CHANGES.md under suiteRevision 15 -- 590 single-site weakenings over 179 vectors, 316 killed, 250 byte-identical, 19 seen and tolerated -- and recorded, in the same entry, that the measurement's full tables are held outside this repository. So the quoted figure could not be re-derived from anything this repository published, and the condition half had never been published at all. It circulated beside a site-axis figure from a later corpus: one method, two corpora, two dates, and the appearance of a disagreement that was never there.

It is reproducible now. docs/PRIOR-FORCING.json pins the earliest forcing baseline this repository carries (c0191622375f, 590 sites) and the tree the original campaign recorded as the one it ran against (6e1a7c388bb2, 179 vectors), and --verify-prior re-derives the projection from those two objects: 54 forced, 24 weak, 0 unforced across 78 conditions, which is the quoted figure exactly. Every one of the seven numbers this page used to carry as a typed constant survives the reconstruction unchanged. That is the finding, not a vindication: a figure nobody could re-derive was indistinguishable from a wrong one until somebody re-derived it, and seven typed constants would have read exactly the same on the day they stopped being true.

The reconstruction is an approximation, and here is its shape. Restricting the later baseline's killer sets to the 179 vectors of the earlier corpus says which weakenings that corpus still catches; it says nothing about whether a weakening it stops catching was tolerated or ignored then. The condition projection never reads that distinction, so the approximation is exact on this axis and silent on the other. 12 recorded weakenings are killed only by vectors that did not exist yet and are dropped; 319 survive, which is 3 more than the released note records for the same corpus.

That residue is not a mystery, and it is not the obvious explanation either. The reference rail is byte-identical at the two pinned commits, so the two runs enumerated the same sites and a moved rail explains nothing. What moved is the instrument. Exactly 3 of the weakenings this reconstruction counts sit on the branch a verifier takes when the consumer supplies no substrate key policy at all:

  • tier.go::(recv).DeriveTiers::IF_DISJ::ded23dc7ada6
  • tier.go::(recv).DeriveTiers::IF_OFF::55f10b0ced17
  • verify.go::anchorPolicyCodes::IF_OFF::ed58b9be9c83

Removing any of those 3 changes nothing on a run that supplies a key policy, because the branch is not taken. Only a run that also makes a no-key pass can see them, and the run behind the released note made no such pass: before the AEE_SUBSTRATE_KEYS channel existed the harness recorded tiers_without_key: None for every run and the evaluator skips a tier column it was handed nothing for. This repository's README and packaging/run_vectors.py both say so, in the section on why each vector is now run twice.

So the whole difference between the two runs closes, and closes exactly: 331 killed in the pinned baseline = 316 in the released note + 12 killed only by vectors added afterwards + 3 the earlier instrument could not observe. --verify-prior derives that 3-site set from the baseline by the policy == nil branch rather than reading this list, refuses if the derived set is not the set named in docs/PRIOR-FORCING.json, refuses if the subtraction does not land on the note's figure, and refuses if the note's four classes do not account for every site in the baseline -- which is what stops the sum closing on two errors that cancel.

What that does not establish, stated because the arithmetic is seductive. The earlier campaign's own per-site table is not in this repository, so nothing here checks, site by site, that those 3 are the ones it scored differently. What is checked is that they are the only candidates the mechanism admits and that the identity lands on the published figure with no slack. An exact closure over a derived set is strong and it is not a per-site diff, and the difference between those two is the sort of thing this page exists to keep visible.

The current figure below is not a constant at all. It is derived from the committed baseline on every run, so the two can never again be quoted from different dates.

The conditions nothing uniquely attributes to their own vectors

Each of these conditions is cited by live vectors that catch real weakenings of the rail. What none of them has is a weakening that its vectors alone catch. The practical reading: a third-party verifier could get the rule in this row wrong in some way this corpus does not separate from a neighbouring rule, and still pass. The kills it shares column is how many weakenings its vectors catch alongside somebody else's.

conditionspec anchorrulevectorskills it shares
aee-c-3L440-442a row carrying a label from the carried caught set contributes fail2105
aee-c-7L447-450UNRESOLVED -- ok-002 is the sole carrier and the corpus does not separate this id from aee-c-2. Candidate reading, recorded rather than asserted: the third recompute condition, which contributes pass_indirect when some clean row is not (substrate, intercepted) and pass when none is190
aee-c-14L557-560clean intercepted row refs arming AND covering sealed5133
aee-c-15L958-960one run-level arming/sealed/examination record covers every row earned under it195
aee-c-16L953-958observationSelectors is producer vocabulary positionally parallel to observationRefs; no gate reads it190
aee-c-25L1744-1747RFC 6962 domain-separated hashing156
aee-c-26L1747-1749RFC 6962 recursive split, never duplicate-pad5116
aee-c-27L1749leaves in array order156
aee-c-28L1749a single-record tree's root is its leaf hash192
aee-c-30L1754-1756batchRoot must recompute370
aee-c-32L1744-1748batchRoot is over every carried record in array order, referenced by a row or not2100
aee-c-33L766-775the evidence tier is derived per row and never carried: artifact is declared, substrate is attested when every covering signature verifies under consumer policy and unattested otherwise, and the tier never alters result1109
aee-c-34L772-774no TOFU: a consumer with no policy-pinned substrate root treats every substrate row as unattested and MUST NOT infer the root from the predicate195
aee-c-35L1901-1903keyid is an unauthenticated lookup hint, never the check195
aee-c-36L1302-1304; L545-546a record signature is DSSE PAE over (payloadType, payload); the byte-pure validity gate never reads a signature, so a signature that does not verify is a tier fact and not a validity fault195
aee-c-38L779-781a carried predicate-level evidenceTier member MUST be ignored190
aee-c-41L992-993basis required, closed {substrate, artifact}1109
aee-c-45L1046-1052weakest-input method composition5120
aee-c-49L1283-1286the literal none is valid on a caught row too, and states that the event was observed and no enforcement layer acted191
aee-c-50L1269-1270actualLayer names the enforcement layer that acted on the row's containment event192
aee-c-61L779-781a predicate-level member beginning with the reserved aee prefix MUST be ignored190
aee-c-62L229-237binding is anti-splice157
aee-c-64L1330-1335sealed record required members599
aee-c-68L1187-1188each referenced record independently satisfies its class constraints299
aee-c-71L1687-1691unknown aeeKind covers nothing3102
aee-c-73L1693-1695the aee payload member prefix is reserved; every other payload member is producer territory and does not stop a record covering5110
aee-c-81L920row attackId appears in the manifest110

The conditions the corpus does not force at all

None. Every condition this corpus cites is caught by at least one of its own vectors. That is the one place this measurement comes out better than the headline suggests, and it is stated with the same weight as the rest.

The failure codes a rail can decline to emit

A separate question, and a harder one: can a conforming verifier decline to emit a failure code at all? It is measured with the one mutation operator that removes exactly one emission and nothing else, because the operator that removes the enclosing statement takes a field parse down with it on a parse block and is then killed for a reason that has nothing to do with the code inside. In the campaign recorded above, removing any single emission below left every vector passing, so nothing in this suite obliges an implementer to produce these codes.

codesites, by outcome
corpus-anchor-mismatch1 DEAD
record-undecodable2 SILENT
result-vocabulary1 DEAD, 1 SILENT
row-attack-unknown1 SILENT
substrate-anchor-mismatch1 DEAD

corpus-anchor-mismatch and substrate-anchor-mismatch are the consumer-policy anchor comparison, which vectors/coverage-unforced.json cell U5 already records as unreachable by a self-contained corpus by construction: the anchor is an operand each consumer supplies, and no vector can carry one. The rest are unminted rather than unmintable, and each is a vector somebody could write.

That is measured over the 45 codes an emission-removal site names directly. 20 of the rail's declared codes are reached only through a shared helper that takes the code as an argument, and 14 emission-removal sites name no code for the same reason. Those codes are not measured on this axis, and they are not counted as forced either: arming-covers-nothing, assessed-set-exceeds-declaration, attribution-pin-unmatched, attribution-pinned-recordless, attribution-unpinnable, caught-row-uncovered, clean-row-contradicted, clean-row-uncovered, examination-covers-nothing, interception-record-orphaned, moat-drop-covers-nothing, observed-attack-uncaught, observed-set-mismatch, payload-commitment-malformed, reconstructed-row-uncovered, record-kind-unknown-covers-nothing, result-recompute-mismatch, sealed-covers-nothing, sealed-record-absent, uncommitted-observation-covers-nothing.

The requirements no vector pins, which is a separate register

A third register, kept separately because it answers a different question. The condition axis above asks what a vector forces; vectors/coverage-unforced.json asks, requirement by requirement, whether a self-contained vector could pin it at all, and docs/COVERAGE-MATRIX.md carries the rows. The counts are here so this page is the whole answer and the matrix stays the detail.

classificationcount
consumer-policy-unvectorable5
forcible-but-unforced1
producer-obligation-ungated1

The sites no vector can ever kill, and why that is not a gap

Two kinds of site are recorded in the baseline as unkillable rather than unforced, because writing a vector for them is wasted work. unkillable is a true equivalent mutant: the weakened rail computes exactly what the original computes, on every input. masked is a real rule that an earlier check reaches first on every input that could get to it, so it cannot be forced as the corpus and the rail stand. Both are claims about the world and the forcing gate falsifies them: an annotated site that is ever killed fails the build.

siteannotationwhy no vector can kill it
recompute.go::Recompute::CASE_DEL::23e37ac1a426unkillableA true equivalent mutant. The arm has an EMPTY body and the switch has no default, so deleting it leaves a value that matches nothing and does nothing -- byte-identical behaviour on every input. No vector can ever kill it and none should be written.
recompute.go::resultRank::CASE_DEL::e24f0559dc93unkillableA true equivalent mutant. Deleting the arm sends ResultFail to default: return -1, and minResult compares ranks only relatively (resultRank(b) < resultRank(a)), so fail still sorts below every known token and every composition reaches the same result. Contrast the sibling case ResultDegraded: return 1, which is KILLED: deleting THAT one drops degraded below fail and a mixed run reports the softer verdict.
validity.go::sealedConstraintsMet::IF_DISJ::1ed486c5f788maskedUnmintable rather than unminted. Removing this disjunct lets a sealed record with no aeePostureDigest reach the posture != pinnedPosture equality in the same branch with posture == "", which refuses it and returns the same recordEval the guard would have. A statement that could tell the two apart needs an EMPTY pinned posture digest, and GATE 0 refuses an environment with no networkPosture (environment-incomplete). bad-903-sealed-missing-posture was built for this rule, mutation-checked, and does not kill. Writing another vector for it is wasted work until the rail or the corpus shape changes; this row exists so the next reader is not sent to do it. The site moved out of evaluateKind when the per-kind constraints were extracted into their own functions, and the move cost this key its positional suffix: the two #N occurrences were the arming branch and the sealed branch of one switch, and they are now one occurrence each in two functions. The claim is unchanged; it is the sealed one.
validity.go::sealedConstraintsMet::IF_DISJ::b65a453ecc06unkillableA true equivalent mutant, by objBool's contract. objBool returns (false, false) for an absent member and for a member that is not a boolean, so !hasStillArmed implies !stillArmed, and `!hasStillArmed

What this measurement does not establish

  • It is a property of a pair. Forcing is measured against one reference rail. A rule this corpus does not force on that rail might still be forced on another whose control flow differs, and the reverse.
  • A weakening operator is not an attacker. The campaign removes guards, disjuncts, conjuncts, emissions and boolean returns one at a time, and restricts one collection loop at a time to a single member. A verifier wrong in a way no single-site weakening expresses is not measured here.
  • The quantifier operator asks about the SECOND member, and cuts both ways. LOOP_FIRST closes a loop body with a break, so a rule the specification states over every member of a carried collection is applied to one member only. Where the loop is a universal, that is a weakening, and a vector that still passes was never forcing the "every": it carried one witness, or its defective member was the one the weakened rail still looks at. Where the loop is instead a SEARCH for one satisfying member, the same edit narrows the search and makes the rail stricter rather than laxer. Both are scored the same way, by replaying the corpus, so a row here says the corpus notices the loop being cut short; read the direction off the site.
  • The condition classes cannot be compared across revisions. See the definition above: attribution depends on the rest of the corpus, so the classification is only meaningful against the inputs named in the provenance table.
  • Some codes are not measured at all, and are listed above rather than assumed either way.
  • The unmeasurable sites stay unmeasurable. The baseline's INCONCLUSIVE class is mutants that did not generate, build, terminate or run cleanly. Nothing is asserted about them and they are never folded into a gap count.

Regenerating every number above

# the document, from the committed data; --check fails instead of writing
python3 scripts/condition-forcing-gate.py --check

# the earlier figure, re-derived from the two git objects it is pinned to
python3 scripts/condition-forcing-gate.py --verify-prior   # needs a full clone

# the same figures under a SECOND implementation that shares no code with the
# generator, because a number confirmed by the program that wrote it is not
# confirmed. It recomputes from the two inputs and reads this page's prose back
python3 scripts/condition-forcing-crosscheck.py --with-prior

# the rail at those two commits, which is why a moved rail explains nothing
git diff --stat 6e1a7c388bb2 c0191622375f -- aee/

# the data itself: re-run the campaign that produces docs/FORCING-BASELINE.json
python3 scripts/forcing-gate.py --scope all --sync   # needs a Go toolchain

# the input digests in the provenance table
sha256sum vectors/MANIFEST.json docs/FORCING-BASELINE.json

The first command reads only files in this repository and needs nothing installed. It is the one to run when checking a number quoted from here. The second, third and fourth read history, so a shallow clone makes them stop rather than pass.

The first command proves only that this page has not drifted from the data, which is a weaker property than it sounds: the generator regenerates the page and compares, so a generator that read the definitions wrongly would produce a wrong page and then agree with it. The crosscheck is the answer to that. It implements the three definitions above independently, recomputes from vectors/MANIFEST.json and docs/FORCING-BASELINE.json, and then reads the headline sentence, the provenance rows and every weak row back out of this page and refuses on any disagreement.