Where this suite is weak
September 2, 2026 ยท View on GitHub
A conformance suite that publishes only what it covers is asking to be taken on trust. This document is the other half: the conditions this corpus cites and does not uniquely force, the failure codes a conforming verifier can decline to emit entirely, and the sites no vector can ever reach. Every number here is derived from the published mutation baseline on every run, by the gate named in the comment above, so it cannot drift from the data it summarises.
The figure, and what it measures
Of the 93 normative conditions this corpus cites, 66 are forced by a vector no other condition's vectors duplicate, 27 are covered only redundantly, and 0 are not forced at all.
| what | value |
|---|---|
| suite | adversarial-execution-evidence-conformance |
| corpus | suiteRevision 28, 272 vectors (61 accept, 209 reject, 2 indeterminate) |
| vendored specification | commit 0dbe10bcc959b63dc42370a5db09812c9476f59a, fetchable from astrogilda/attestation at branch predicate/adversarial-execution-evidence |
| reviewed at | in-toto/attestation#570 |
vectors/MANIFEST.json | sha256:7b6c7cc5f0068f5070f0da1eef2d43387eb070858cfd4041f850674de29471f0 |
docs/FORCING-BASELINE.json | sha256:8a108f068cb98418661a9a84b4b94dbf9655a99d8cb4be7c4f537cae812de5fb |
| campaign | 807 single-site weakenings: 459 KILLED, 32 SILENT, 311 DEAD, 5 INCONCLUSIVE |
Every vector in the corpus is depended on by at least one recorded weakening, so no vector is wholly redundant and the campaign covers the corpus as it stands.
The corpus cites the specification's rules by aee-c-NN condition id on every
vector. Take the vectors citing one condition, and ask what the mutation campaign
in docs/FORCING-BASELINE.json says about them:
- forced -- some rail weakening is caught by those vectors and by no others. Delete them and that weakening goes undetected.
- weak -- they catch weakenings, but every one is also caught by a vector citing something else. The rule is covered; nothing is uniquely attributable to this condition.
- not forced -- they catch no weakening at all.
Weak is a statement about attribution, not about strength, and it does not move in one direction. Adding a vector that happens to catch an already-caught weakening enlarges that weakening's killer set, which can move a condition from forced to weak while the corpus got strictly better. Two of these figures from different revisions are not a trend line, and a rise in the weak count is not a regression.
The figure this replaces, and what could not be re-derived about it. A condition
projection was quoted before this page existed: 24 of 78 conditions weak, over the
corpus at suiteRevision 14. Its own campaign put its site-axis half in
vectors/CHANGES.md under suiteRevision 15 -- 590 single-site weakenings over 179
vectors, 316 killed, 250 byte-identical, 19 seen and tolerated -- and recorded, in the
same entry, that the measurement's full tables are held outside this repository. So
the quoted figure could not be re-derived from anything this repository published, and
the condition half had never been published at all. It circulated beside a site-axis
figure from a later corpus: one method, two corpora, two dates, and the appearance of a
disagreement that was never there.
It is reproducible now. docs/PRIOR-FORCING.json pins the earliest forcing baseline
this repository carries (c0191622375f, 590 sites) and the tree the original campaign
recorded as the one it ran against (6e1a7c388bb2, 179 vectors), and --verify-prior
re-derives the projection from those two objects: 54 forced, 24 weak, 0 unforced
across 78 conditions, which is the quoted figure exactly. Every one of the seven
numbers this page used to carry as a typed constant survives the reconstruction
unchanged. That is the finding, not a vindication: a figure nobody could re-derive was
indistinguishable from a wrong one until somebody re-derived it, and seven typed
constants would have read exactly the same on the day they stopped being true.
The reconstruction is an approximation, and here is its shape. Restricting the later baseline's killer sets to the 179 vectors of the earlier corpus says which weakenings that corpus still catches; it says nothing about whether a weakening it stops catching was tolerated or ignored then. The condition projection never reads that distinction, so the approximation is exact on this axis and silent on the other. 12 recorded weakenings are killed only by vectors that did not exist yet and are dropped; 319 survive, which is 3 more than the released note records for the same corpus.
That residue is not a mystery, and it is not the obvious explanation either. The reference rail is byte-identical at the two pinned commits, so the two runs enumerated the same sites and a moved rail explains nothing. What moved is the instrument. Exactly 3 of the weakenings this reconstruction counts sit on the branch a verifier takes when the consumer supplies no substrate key policy at all:
tier.go::(recv).DeriveTiers::IF_DISJ::ded23dc7ada6tier.go::(recv).DeriveTiers::IF_OFF::55f10b0ced17verify.go::anchorPolicyCodes::IF_OFF::ed58b9be9c83
Removing any of those 3 changes nothing on a run that supplies a key policy, because the
branch is not taken. Only a run that also makes a no-key pass can see them, and the run
behind the released note made no such pass: before the AEE_SUBSTRATE_KEYS channel
existed the harness recorded tiers_without_key: None for every run and the evaluator
skips a tier column it was handed nothing for. This repository's README and
packaging/run_vectors.py both say so, in the section on why each vector is now run
twice.
So the whole difference between the two runs closes, and closes exactly: 331 killed in
the pinned baseline = 316 in the released note + 12 killed only by vectors added
afterwards + 3 the earlier instrument could not observe. --verify-prior derives that
3-site set from the baseline by the policy == nil branch rather than reading this
list, refuses if the derived set is not the set named in docs/PRIOR-FORCING.json,
refuses if the subtraction does not land on the note's figure, and refuses if the note's
four classes do not account for every site in the baseline -- which is what stops the
sum closing on two errors that cancel.
What that does not establish, stated because the arithmetic is seductive. The earlier campaign's own per-site table is not in this repository, so nothing here checks, site by site, that those 3 are the ones it scored differently. What is checked is that they are the only candidates the mechanism admits and that the identity lands on the published figure with no slack. An exact closure over a derived set is strong and it is not a per-site diff, and the difference between those two is the sort of thing this page exists to keep visible.
The current figure below is not a constant at all. It is derived from the committed baseline on every run, so the two can never again be quoted from different dates.
The conditions nothing uniquely attributes to their own vectors
Each of these conditions is cited by live vectors that catch real weakenings of the
rail. What none of them has is a weakening that its vectors alone catch. The practical
reading: a third-party verifier could get the rule in this row wrong in some way this
corpus does not separate from a neighbouring rule, and still pass. The kills it shares
column is how many weakenings its vectors catch alongside somebody else's.
| condition | spec anchor | rule | vectors | kills it shares |
|---|---|---|---|---|
aee-c-3 | L440-442 | a row carrying a label from the carried caught set contributes fail | 2 | 105 |
aee-c-7 | L447-450 | UNRESOLVED -- ok-002 is the sole carrier and the corpus does not separate this id from aee-c-2. Candidate reading, recorded rather than asserted: the third recompute condition, which contributes pass_indirect when some clean row is not (substrate, intercepted) and pass when none is | 1 | 90 |
aee-c-14 | L557-560 | clean intercepted row refs arming AND covering sealed | 5 | 133 |
aee-c-15 | L958-960 | one run-level arming/sealed/examination record covers every row earned under it | 1 | 95 |
aee-c-16 | L953-958 | observationSelectors is producer vocabulary positionally parallel to observationRefs; no gate reads it | 1 | 90 |
aee-c-25 | L1744-1747 | RFC 6962 domain-separated hashing | 1 | 56 |
aee-c-26 | L1747-1749 | RFC 6962 recursive split, never duplicate-pad | 5 | 116 |
aee-c-27 | L1749 | leaves in array order | 1 | 56 |
aee-c-28 | L1749 | a single-record tree's root is its leaf hash | 1 | 92 |
aee-c-30 | L1754-1756 | batchRoot must recompute | 3 | 70 |
aee-c-32 | L1744-1748 | batchRoot is over every carried record in array order, referenced by a row or not | 2 | 100 |
aee-c-33 | L766-775 | the evidence tier is derived per row and never carried: artifact is declared, substrate is attested when every covering signature verifies under consumer policy and unattested otherwise, and the tier never alters result | 1 | 109 |
aee-c-34 | L772-774 | no TOFU: a consumer with no policy-pinned substrate root treats every substrate row as unattested and MUST NOT infer the root from the predicate | 1 | 95 |
aee-c-35 | L1901-1903 | keyid is an unauthenticated lookup hint, never the check | 1 | 95 |
aee-c-36 | L1302-1304; L545-546 | a record signature is DSSE PAE over (payloadType, payload); the byte-pure validity gate never reads a signature, so a signature that does not verify is a tier fact and not a validity fault | 1 | 95 |
aee-c-38 | L779-781 | a carried predicate-level evidenceTier member MUST be ignored | 1 | 90 |
aee-c-41 | L992-993 | basis required, closed {substrate, artifact} | 1 | 109 |
aee-c-45 | L1046-1052 | weakest-input method composition | 5 | 120 |
aee-c-49 | L1283-1286 | the literal none is valid on a caught row too, and states that the event was observed and no enforcement layer acted | 1 | 91 |
aee-c-50 | L1269-1270 | actualLayer names the enforcement layer that acted on the row's containment event | 1 | 92 |
aee-c-61 | L779-781 | a predicate-level member beginning with the reserved aee prefix MUST be ignored | 1 | 90 |
aee-c-62 | L229-237 | binding is anti-splice | 1 | 57 |
aee-c-64 | L1330-1335 | sealed record required members | 5 | 99 |
aee-c-68 | L1187-1188 | each referenced record independently satisfies its class constraints | 2 | 99 |
aee-c-71 | L1687-1691 | unknown aeeKind covers nothing | 3 | 102 |
aee-c-73 | L1693-1695 | the aee payload member prefix is reserved; every other payload member is producer territory and does not stop a record covering | 5 | 110 |
aee-c-81 | L920 | row attackId appears in the manifest | 1 | 10 |
The conditions the corpus does not force at all
None. Every condition this corpus cites is caught by at least one of its own vectors. That is the one place this measurement comes out better than the headline suggests, and it is stated with the same weight as the rest.
The failure codes a rail can decline to emit
A separate question, and a harder one: can a conforming verifier decline to emit a failure code at all? It is measured with the one mutation operator that removes exactly one emission and nothing else, because the operator that removes the enclosing statement takes a field parse down with it on a parse block and is then killed for a reason that has nothing to do with the code inside. In the campaign recorded above, removing any single emission below left every vector passing, so nothing in this suite obliges an implementer to produce these codes.
| code | sites, by outcome |
|---|---|
corpus-anchor-mismatch | 1 DEAD |
record-undecodable | 2 SILENT |
result-vocabulary | 1 DEAD, 1 SILENT |
row-attack-unknown | 1 SILENT |
substrate-anchor-mismatch | 1 DEAD |
corpus-anchor-mismatch and substrate-anchor-mismatch are the consumer-policy anchor
comparison, which vectors/coverage-unforced.json cell U5 already records as
unreachable by a self-contained corpus by construction: the anchor is an operand each
consumer supplies, and no vector can carry one. The rest are unminted rather than
unmintable, and each is a vector somebody could write.
That is measured over the 45 codes an emission-removal site names directly. 20 of the
rail's declared codes are reached only through a shared helper that takes the code as an
argument, and 14 emission-removal sites name no code for the same reason. Those codes
are not measured on this axis, and they are not counted as forced either:
arming-covers-nothing, assessed-set-exceeds-declaration,
attribution-pin-unmatched, attribution-pinned-recordless, attribution-unpinnable,
caught-row-uncovered, clean-row-contradicted, clean-row-uncovered,
examination-covers-nothing, interception-record-orphaned,
moat-drop-covers-nothing, observed-attack-uncaught, observed-set-mismatch,
payload-commitment-malformed, reconstructed-row-uncovered,
record-kind-unknown-covers-nothing, result-recompute-mismatch,
sealed-covers-nothing, sealed-record-absent,
uncommitted-observation-covers-nothing.
The requirements no vector pins, which is a separate register
A third register, kept separately because it answers a different question. The condition
axis above asks what a vector forces; vectors/coverage-unforced.json asks, requirement
by requirement, whether a self-contained vector could pin it at all, and
docs/COVERAGE-MATRIX.md carries the rows. The counts are here so this page is the
whole answer and the matrix stays the detail.
| classification | count |
|---|---|
| consumer-policy-unvectorable | 5 |
| forcible-but-unforced | 1 |
| producer-obligation-ungated | 1 |
The sites no vector can ever kill, and why that is not a gap
Two kinds of site are recorded in the baseline as unkillable rather than unforced,
because writing a vector for them is wasted work. unkillable is a true equivalent
mutant: the weakened rail computes exactly what the original computes, on every input.
masked is a real rule that an earlier check reaches first on every input that could
get to it, so it cannot be forced as the corpus and the rail stand. Both are claims
about the world and the forcing gate falsifies them: an annotated site that is ever
killed fails the build.
| site | annotation | why no vector can kill it |
|---|---|---|
recompute.go::Recompute::CASE_DEL::23e37ac1a426 | unkillable | A true equivalent mutant. The arm has an EMPTY body and the switch has no default, so deleting it leaves a value that matches nothing and does nothing -- byte-identical behaviour on every input. No vector can ever kill it and none should be written. |
recompute.go::resultRank::CASE_DEL::e24f0559dc93 | unkillable | A true equivalent mutant. Deleting the arm sends ResultFail to default: return -1, and minResult compares ranks only relatively (resultRank(b) < resultRank(a)), so fail still sorts below every known token and every composition reaches the same result. Contrast the sibling case ResultDegraded: return 1, which is KILLED: deleting THAT one drops degraded below fail and a mixed run reports the softer verdict. |
validity.go::sealedConstraintsMet::IF_DISJ::1ed486c5f788 | masked | Unmintable rather than unminted. Removing this disjunct lets a sealed record with no aeePostureDigest reach the posture != pinnedPosture equality in the same branch with posture == "", which refuses it and returns the same recordEval the guard would have. A statement that could tell the two apart needs an EMPTY pinned posture digest, and GATE 0 refuses an environment with no networkPosture (environment-incomplete). bad-903-sealed-missing-posture was built for this rule, mutation-checked, and does not kill. Writing another vector for it is wasted work until the rail or the corpus shape changes; this row exists so the next reader is not sent to do it. The site moved out of evaluateKind when the per-kind constraints were extracted into their own functions, and the move cost this key its positional suffix: the two #N occurrences were the arming branch and the sealed branch of one switch, and they are now one occurrence each in two functions. The claim is unchanged; it is the sealed one. |
validity.go::sealedConstraintsMet::IF_DISJ::b65a453ecc06 | unkillable | A true equivalent mutant, by objBool's contract. objBool returns (false, false) for an absent member and for a member that is not a boolean, so !hasStillArmed implies !stillArmed, and `!hasStillArmed |
What this measurement does not establish
- It is a property of a pair. Forcing is measured against one reference rail. A rule this corpus does not force on that rail might still be forced on another whose control flow differs, and the reverse.
- A weakening operator is not an attacker. The campaign removes guards, disjuncts, conjuncts, emissions and boolean returns one at a time, and restricts one collection loop at a time to a single member. A verifier wrong in a way no single-site weakening expresses is not measured here.
- The quantifier operator asks about the SECOND member, and cuts both
ways.
LOOP_FIRSTcloses a loop body with abreak, so a rule the specification states over every member of a carried collection is applied to one member only. Where the loop is a universal, that is a weakening, and a vector that still passes was never forcing the "every": it carried one witness, or its defective member was the one the weakened rail still looks at. Where the loop is instead a SEARCH for one satisfying member, the same edit narrows the search and makes the rail stricter rather than laxer. Both are scored the same way, by replaying the corpus, so a row here says the corpus notices the loop being cut short; read the direction off the site. - The condition classes cannot be compared across revisions. See the definition above: attribution depends on the rest of the corpus, so the classification is only meaningful against the inputs named in the provenance table.
- Some codes are not measured at all, and are listed above rather than assumed either way.
- The unmeasurable sites stay unmeasurable. The baseline's INCONCLUSIVE class is mutants that did not generate, build, terminate or run cleanly. Nothing is asserted about them and they are never folded into a gap count.
Regenerating every number above
# the document, from the committed data; --check fails instead of writing
python3 scripts/condition-forcing-gate.py --check
# the earlier figure, re-derived from the two git objects it is pinned to
python3 scripts/condition-forcing-gate.py --verify-prior # needs a full clone
# the same figures under a SECOND implementation that shares no code with the
# generator, because a number confirmed by the program that wrote it is not
# confirmed. It recomputes from the two inputs and reads this page's prose back
python3 scripts/condition-forcing-crosscheck.py --with-prior
# the rail at those two commits, which is why a moved rail explains nothing
git diff --stat 6e1a7c388bb2 c0191622375f -- aee/
# the data itself: re-run the campaign that produces docs/FORCING-BASELINE.json
python3 scripts/forcing-gate.py --scope all --sync # needs a Go toolchain
# the input digests in the provenance table
sha256sum vectors/MANIFEST.json docs/FORCING-BASELINE.json
The first command reads only files in this repository and needs nothing installed. It is the one to run when checking a number quoted from here. The second, third and fourth read history, so a shallow clone makes them stop rather than pass.
The first command proves only that this page has not drifted from the data,
which is a weaker property than it sounds: the generator regenerates the page
and compares, so a generator that read the definitions wrongly would produce a
wrong page and then agree with it. The crosscheck is the answer to that. It
implements the three definitions above independently, recomputes from
vectors/MANIFEST.json and docs/FORCING-BASELINE.json, and then reads the
headline sentence, the provenance rows and every weak row back out of this page
and refuses on any disagreement.