Documentation Index
September 17, 2026 · View on GitHub
This index covers the current public contracts, operator guides, and reproducible research context for LegalForecastBench. The official-run runbook and reproduction guide are checked against the CLI by automated tests. Corrections are welcome as issues or pull requests.
Start Here
| If you want to… | Read |
|---|---|
| Understand what the benchmark measures and how | METHODS.md |
| Reproduce or audit a published result | reproduce-or-audit.md |
| Know what may and may not be claimed publicly | publication-governance.md |
| Operate a protected official cycle | official-run-runbook.md |
| Submit a community harness comparison | community-contributor-guide.md, then multiharness-adapter-spec.md and community-submissions.md |
Official Benchmark
-
METHODS.md: eval-card-grade methods — construct, frozen inputs, leakage controls, metrics, inference, related work, human-baseline status, limitations, and withdrawal policy.
-
Contamination-tier reporting: the mechanical rule that distinguishes contamination-resistant scores from preliminary (non-contamination-resistant) scores — both official, viable results — and the paired drift metric between them.
-
official-run-runbook.md: the public release boundary — immutable inputs, protected forecast/fan-in workflows, strict scoring, reporting, and hold conditions.
-
reproduce-or-audit.md: credential-free reproduction of public arithmetic and the deeper audit workflow.
-
Publication governance: current track-separation, arm, reporting, and non-affiliation rules.
-
Hugging Face benchmark publication: manually gated access, immutable dataset revisions, native leaderboard registration, and short-lived automated publication.
-
Jev one-shot condition: unchanged prediction units with full text or persisted Luna document summaries.
Corpus Handoff Boundary
Corpus construction, private source bytes, selection, unitization, and quality control are owned by the companion LegalForecastCorpus repository. This public repository receives only immutable, outcome-blinded release inputs through the documented public/private release boundary.
- Commitment contracts: named canonical-byte, digest-representation, and schema-domain APIs for retained public release code.
Community Multi-Harness (non-official)
The multi-harness layer is a separate, non-official track. Its results never rank alongside official results.
- community-contributor-guide.md: install, probe, select, run, interrupt, resume, validate, and package from a dedicated environment.
- multiharness-adapter-spec.md: the community adapter contract.
- community-submissions.md: submission packaging, attestations, credits, funding policy, and PR intake.
- harness-efficiency-observations.md: receipt-backed duration, cost, and token accounting published as peer columns.
- Multiharness contracts: the current sealed-deliverable, evaluation, scoring, identity, compatibility, and Harvey LAB source boundaries.
- multiharness-receipt-authority.md: the external Ed25519 evaluator-issuer seam and credential-free executable probe procedure.
Adapter tracks
- provider-baselines.md: provider/runtime reference points and the publication terms they rest on.
- Local CLI adapter manifest: the current generic manifest contract for local agentic CLI adapters.
Public Release Schema Reference
Only the outcome-blinded forecast release, separately controlled labels release, and public/private release boundary below are active Bench contracts.
- Public/private release boundary: the additive split between outcome-blinded public execution inputs and separately controlled labels.
- Forecast release v1: canonical cases, prediction units, model-visible document indexes, packets, prompts, and byte commitments.
- Labels release v1: separately bound unit outcomes and scoring policy with no forecast-execution API path.
Model metadata
- Model release dates: model anchor sources and dates.