Evaluation
August 21, 2026 ยท View on GitHub
routing-v1 tests the Skill's first job: choose the correct DSH delivery mode, open version-matched authorities, separate Host and Client responsibilities, and select evidence that reaches the real entry path.
Each case runs in a fresh agent context with the Skill and a read-only DSH target. Expected invariants stay hidden from the execution agent. The maintainer scores concrete decisions and cited evidence; preferred wording is irrelevant.
The initial maintainer-recorded receipt is receipts/2026-08-21-routing-v1.json. All four cases were scored as routing and acceptance-plan passes. Raw outputs were not published and no case performed an implementation run, so the receipt is not independently reproducible and does not support claims about generated-code correctness, success-at-one, time-to-green, token savings, or production reliability.
Stronger evidence gate
Before claiming that this Skill improves plugin implementation, run an A/B benchmark against repository instructions alone:
- Pin the same DSH commit, model, reasoning level, permissions, and clean fixture for both arms.
- Keep graders and hard invariants hidden from the execution agent.
- Test real plugin behavior through Loader/app composition, artifact installation, browser interaction, or the live dynamic runtime as the task requires.
- Count a run as successful only when no human patch is needed and every safety invariant passes.
- Publish raw pass/fail counts, time-to-green, agent-request token usage when available, and all unrun or failed cases. Stars and installs remain adoption signals, not substitutes.
Compatibility reviews rerun routing-v1. Implementation-effectiveness claims require a separately versioned, receipt-backed suite.