Evaluation

August 21, 2026 ยท View on GitHub

routing-v1 tests the Skill's first job: choose the correct DSH delivery mode, open version-matched authorities, separate Host and Client responsibilities, and select evidence that reaches the real entry path.

Each case runs in a fresh agent context with the Skill and a read-only DSH target. Expected invariants stay hidden from the execution agent. The maintainer scores concrete decisions and cited evidence; preferred wording is irrelevant.

The initial maintainer-recorded receipt is receipts/2026-08-21-routing-v1.json. All four cases were scored as routing and acceptance-plan passes. Raw outputs were not published and no case performed an implementation run, so the receipt is not independently reproducible and does not support claims about generated-code correctness, success-at-one, time-to-green, token savings, or production reliability.

Stronger evidence gate

Before claiming that this Skill improves plugin implementation, run an A/B benchmark against repository instructions alone:

  1. Pin the same DSH commit, model, reasoning level, permissions, and clean fixture for both arms.
  2. Keep graders and hard invariants hidden from the execution agent.
  3. Test real plugin behavior through Loader/app composition, artifact installation, browser interaction, or the live dynamic runtime as the task requires.
  4. Count a run as successful only when no human patch is needed and every safety invariant passes.
  5. Publish raw pass/fail counts, time-to-green, agent-request token usage when available, and all unrun or failed cases. Stars and installs remain adoption signals, not substitutes.

Compatibility reviews rerun routing-v1. Implementation-effectiveness claims require a separately versioned, receipt-backed suite.