Legal Benchmark Engine
June 1, 2026 ยท View on GitHub
Status: implemented_early for the Rust schema contract and Harvey compatibility scanner.
Psionic owns the Rust execution and evaluation substrate for legal-agent
benchmark runs. The first landed contract is the schema foundation in
crates/psionic-eval/src/legal_benchmark.rs.
Operating boundary: the upstream Harvey Python harness is reference/backfill only. Owned Harvey-compatible execution runs through Psionic's Rust legal benchmark engine, while Autopilot Blueprint/Program policy selects the upgradable prompts/modules, provider adapter, judge policy, release gates, and promotion rules. Provider names such as Gemini, OpenAI-compatible local servers, or Qwen fine-tunes are adapter metadata, not benchmark authority.
Current Contract
The psionic-eval legal benchmark module defines:
BenchmarkTaskSpecArtifactManifestSourceArtifactDeliverableSpecCriterionSpecJudgePolicyToolPolicyRunConfigRunRecordTranscriptEventToolCallRecordRunMetricsCoverageSnapshotCriterionResultScoreReportComparisonReport
The first fine-tuning export path is now split across:
crates/psionic-data/src/legal_benchmark_training_record.rscrates/psionic-eval/src/legal_benchmark_training_records.rsdocs/LEGAL_BENCHMARK_TRAINING_RECORDS.md
That path exports legal task, run, coverage, transcript, and score artifacts
into legal_benchmark_training_record.v1 bundles for Qwen-family adapter
smoke work. The exporter keeps judge-only scoring data separate from
model-visible examples and excludes model-visible examples when a run exposed
hidden criteria.
Every top-level contract has an explicit schema version. The run contract requires task identity, task version, input artifact manifest hash, run config hash, and output artifact manifest hash, so later runner work cannot produce a score without the immutable execution identity Autopilot needs.
The Rust agent runner must not add answer text to model output. Run 015 did that through an output-scaffold path and therefore does not count as a benchmark score. Keep it only as a diagnostic example of a bad runner design: the scorer could be fooled when the runner inserted the words the scorer was looking for.
The current supported no-cheat runner metadata is limited to controls that do not rewrite the deliverable:
max_output_tokens: per-run override for the model request output budget.force_write_until_required_deliverables: keep prompting until the model itself writes the required deliverable or the run budget is exhausted.force_validate_after_write: require a model-authored validation step after a model-authored write.plain_text_tool_protocol: ask a weak local model to emit plain JSON tool requests in text, then execute only the JSON tool call the model wrote.
Those controls affect request shape, turn order, and tool-call parsing. They do not add legal analysis, coverage markers, citations, headings, or scoring phrases to output files.
The runner now emits an answer_integrity report in both run receipts and run
metadata. The report records every required or declared answer file, its
pre-score and post-score hash, size, mtime, writer tool-call id, and actor
classification. A scored answer file is valid only when a model-authored
write or edit tool call produced the final bytes. Files created by the
harness, files changed during scoring, missing required files, and files whose
hash no longer matches the model write receipt are marked invalid.
The evaluator reads the same integrity report before and after scoring. If any
answer file changes during scoring, the score report is forced invalid and the
diagnostic is retained in failure_diagnostics. This keeps invalid runs
auditable without allowing them into promotion metrics.
crates/psionic-eval/src/legal_benchmark_schema.rs adds the canonical v1
schema layer used for long-lived legal benchmark receipts and training data.
It wraps today's RunRecord, output manifest, score report, and
answer_integrity report into a LegalRunReceipt with:
- benchmark id and visibility
- base model, adapter, tokenizer, prompt template, thinking mode, and tools
- source document or input-manifest hashes
- transcript action hashes and tool-call hashes
- answer file hashes and actor information
- scorer version, score hash, wall-clock timings, git commit, dirty-tree flag, worker id, hardware summary, replay command, and artifact refs
The schema module also defines the first Rust shapes for legal training
examples, bad-run examples, preference pairs, reward traces, dataset
manifests, adapter manifests, model candidates, promotion decisions, Pylon
training jobs, Pylon worker receipts, Psionic training configs, and Psionic
training receipts. These are plain canonical JSON contracts for now; moving
them into a shared crate is deferred until psionic-train needs to consume
the same stable structs directly.
Validation rejects receipts that omit benchmark visibility, answer file content hashes, scorer version, replay command, or required artifact hashes. Use:
cargo test -p psionic-eval legal_benchmark_schema
cargo run -p psionic-eval --example legal_benchmark_validate_run_receipt -- <path>
cargo run -p psionic-eval --example legal_benchmark_print_run_summary -- <path>
crates/psionic-eval/src/legal_benchmark_failed_trajectory.rs captures failed
legal runs as complete LegalBadRunExample artifacts. A bad-run example keeps
the full prompt, full model response, tool-call transcript, attempted writes,
required file status, answer content and hashes when present, action sequence,
stop reason, score, scorer feedback, integrity status, failure class,
suggested correction, and training eligibility. Bad examples are never marked
SFT-eligible directly; a later dataset builder has to convert them into a
positive example or preference pair.
Failure capture currently classifies missing files, wrong output paths, empty or badly sized answers, missing source use, hallucinated citations, malformed or invalid JSON tool calls, harness integrity failures, scorer outages, timeouts, refusals, missing submissions, and uncategorized failures. Hidden and private benchmark failures stay audit-only unless an explicit private-training override is supplied and no hidden labels or scorer secrets are present.
Use:
cargo test -p psionic-eval failed_trajectory_capture
cargo run -p psionic-eval --example legal_benchmark_inspect_failures -- <run-dir>
crates/psionic-eval/src/legal_benchmark_signature_routing.rs adds the
public/synthetic Harvey-compatible signature-routing fixture lane. It maps
structured legal failure families to Probe seed signatures without keyword
matching over prompts and without exposing hidden Harvey labels. The first
fixture suite covers missing deliverables, wrong output paths, missing source
grounding, missing citation provenance, answer-integrity failures, and
judge-supervisor evidence gaps.
The retained routing report lives at
fixtures/legal_benchmark/signature_routing/harvey_public_synthetic_signature_routing_report.json.
Current result: six fixtures, 10000 bps selector pass rate, raw Codex
fixture mean 2222 bps, Probe+Codex fixture mean 10000 bps, and mean delta
7777 bps. This is a deterministic workflow/evidence fixture, not a private
Harvey score or proof of live legal quality.
Use:
cargo test -p psionic-eval --no-default-features --lib legal_benchmark_signature_routing
cargo run -q -p psionic-eval --no-default-features --example legal_benchmark_signature_routing_report
The detailed contract is documented in
docs/LEGAL_BENCHMARK_SIGNATURE_ROUTING.md.
Rust GRPO Smoke Trainer
psionic-train now has a Rust-only legal GRPO smoke command:
cargo run -p psionic-train -- grpo --config configs/legal/qwen36_grpo_smoke.json
What it does today:
- samples deterministic local completion groups for public legal workflow prompts
- scores each completion with a verifier-style reward plugin for file writing, correct path, non-empty answer, source use, submission, and answer integrity
- normalizes reward inside each prompt group
- applies adapter-only weighted updates through the open-adapter backend
- preserves lower-reward completions in
reward_traces.jsonlfor later preference/RL data building - writes
adapter.safetensors,loss_curve.json,checkpoint_summary.json,reward_traces.jsonl, andtraining_receipt.json
Recorded local smoke result:
- run id:
qwen36-legal-grpo-smoke - prompt groups:
2 - sampled completions:
6 - bad completions preserved:
4 - completed steps:
8 - initial file-write preference accuracy:
0.5 - final file-write preference accuracy:
1.0 - initial average reward margin:
-0.0031223297 - final average reward margin:
15.182666 - adapter digest:
825b2d81aeae56d395a4fee7608eead91adf25ac24bae9ff995959df2b95732f - receipt digest:
f030e22b3590c8b5bf51bf355e7eedf84963cf7bccc177499140e91c2edcaf32
The exported adapter also runs through the same Rust legal eval suite:
cargo run -p psionic-eval --example legal_benchmark_eval_suite -- \
--suite suites/harvey_public_three.json \
--model Qwen/Qwen3.6-27B \
--adapter target/legal/qwen36_grpo_smoke/adapter.safetensors \
--out target/legal/qwen36_grpo_eval_smoke
Recorded eval result:
- base score:
3333bps - adapter score:
10000bps - delta:
6667bps - report hash:
df8cbe47739b27de9cd9ca629f51de789cd8584ca8588242ff31b1826efea215
This proves the Rust GRPO training path, reward traces, adapter export, and eval compatibility on a synthetic local smoke. It is not proof of full dense Qwen3.6 RL, distributed Pylon sampling, or performance on private Harvey benchmark tasks.
Pylon Worker Job Protocol
crates/psionic-train/src/qwen_legal_pylon_training_job.rs defines the first
worker job protocol for sending legal benchmark training work to Pylon
workers. It supports dataset shard builds, SFT shard work, DPO shard work,
GRPO sampling, GRPO shard work, eval shards, adapter merges, and artifact
verification.
The job spec records the parent run, model and adapter identity, dataset and training config hashes, shard assignment, expected inputs, required outputs, hardware needs, budget metadata, and receipt requirements. The worker receipt records the worker id, public key, job id, input and output hashes, timings, hardware summary, Psionic version, git commit, logs hash, metrics, failure reason when there is one, Ed25519 signature, and receipt digest.
Run the current local dataset-shard fixture with:
cargo run -p psionic-train --example qwen_legal_pylon_worker_run_once -- \
--job fixtures/qwen_legal/pylon_training_jobs/dataset_shard_job_v1.json
Then verify the receipt with:
cargo run -p psionic-train --example qwen_legal_verify_worker_receipt -- \
target/legal/pylon_jobs/job.qwen-legal.dataset-shard.000001.receipt.json
The matching eval-shard fixture is
fixtures/qwen_legal/pylon_training_jobs/eval_shard_job_v1.json.
This is a local protocol proof. It does not train a live Qwen model by itself; it gives Nexus/Pylon the signed job and receipt shape needed before those workers run real SFT, DPO, or GRPO shards.
Distributed Dataset Sharding
crates/psionic-data/src/legal_benchmark_dataset_sharding.rs owns the first
legal fine-tuning dataset sharding and artifact transport contract for Pylon.
It reads legal_sft_v1 JSONL, requires every row to have a stable
example_id, sorts rows by example_id, computes a canonical global dataset
hash over the sorted rows, and assigns rows to shards with:
shard = sha256(example_id) mod shard_count
The sharder writes:
- one
dataset_shard_manifest.json - one
shard-00000.jsonlstyle file per shard - one uploaded artifact copy per shard under
artifacts/
The manifest records the global dataset hash, source file hash, immutable dataset lock, shard hashes, local artifact refs, uploaded artifact refs, and its own manifest hash. The worker receipt path refuses a shard when the expected manifest hash does not match, verifies the shard artifact before credit, and records a stable credit key so retries do not double-count a shard.
Run the CLI with:
cargo run -p psionic-data --example legal_benchmark_shard_dataset -- \
--dataset fixtures/legal_benchmark/sharding/legal-sft-v1.jsonl \
--shards 4 \
--out target/legal/dataset_shards
The same command shape works for the normal generated dataset path:
cargo run -p psionic-data --example legal_benchmark_shard_dataset -- \
--dataset datasets/legal-sft-v1.jsonl \
--shards 8 \
--out datasets/shards/
crates/psionic-data/src/legal_benchmark_sft_dataset.rs builds canonical
legal_sft_v1 JSONL from honest successful receipts and training-eligible
bad-run examples. It refuses hidden/private-by-default receipts, unknown
visibility, answer-integrity failures, non-model answer mutation, hidden
scoring labels, scorer-only target labels, and known harness-injected marker
runs. Successful runs become golden workflow, source-grounded answer,
tool-discipline, and minimal-answer examples when answer file content is
available. Failed runs stay excluded from raw SFT, but the builder may convert
training-eligible failures into correction examples.
Use:
cargo run -p psionic-data --example legal_benchmark_build_sft_dataset -- \
--runs ./runs \
--out ./datasets/legal-sft-v1.jsonl \
--manifest ./datasets/legal-sft-v1.manifest.json
crates/psionic-data/src/legal_benchmark_dpo_dataset.rs builds canonical
legal_dpo_v1 JSONL preference pairs from honest good runs and
training-eligible bad-run examples. Each pair has the Rust/Psionic-native DPO
shape:
prompt: system message plus the original task prompt and one explicit training focuschosen: the model-written legal answer or correct tool trajectory to imitaterejected: the bad model response or broken trajectory to avoidreason: normalized failure class, such asDidNotWriteRequiredFilesource_run_ids,visibility, andexclusion_flags
The builder emits file-discipline, correct-path, source-grounding, conciseness, submission, and integrity-safe pair families. Its manifest records total pair count, excluded input count, pair counts by failure class, pair counts by family, source receipt refs, excluded input reasons, and a dataset hash.
The DPO builder uses the same safety boundary as the SFT builder: hidden and
private benchmark labels are rejected, unknown visibility is rejected,
integrity-invalid successful runs are rejected, harness-authored answer files
are rejected, hidden/scorer-only markers are rejected, and integrity-invalid or
harness-assisted bad runs are excluded instead of being used as negative
examples. The checked local fixture currently yields 22 public-training pairs
from one good run and one DidNotWriteRequiredFile bad run.
Use:
cargo run -p psionic-data --example legal_benchmark_build_dpo_dataset -- \
--runs ./runs \
--out ./datasets/legal-dpo-v1.jsonl
The optional --manifest ./datasets/legal-dpo-v1.manifest.json flag overrides
the default manifest path derived from the JSONL output path. The loader
function load_legal_dpo_dataset is the handoff point for the Psionic DPO
trainer.
Qwen3.6 prompt handling is Rust-native. crates/psionic-models/src/qwen36.rs
renders Qwen3.6 chat prompts in explicit Thinking, DirectAnswer, and
MixedExplicit modes without using /think or /nothink soft-switch tokens.
The renderer supports tokenizer JSON loading through the Rust tokenizers
crate, tool-response transcript rendering, optional empty think-block emission
for direct-answer generation prompts, deterministic prompt hashes, and a small
Qwen36PromptReceipt that can be embedded in later run receipts.
crates/psionic-transformer/src/qwen36_loss_masks.rs owns the matching loss
mask contract: system, user, and tool spans are ignored under assistant-only
loss, and empty think blocks can be ignored explicitly.
The SFT dataset example schema now also carries reasoning_mode, so legal
examples can declare direct_answer or a later thinking-mode value without
using hidden prompt switches.
Use:
cargo test -p psionic-models qwen36_template
cargo test -p psionic-transformer qwen36_loss_masks
The first Rust-only Qwen3.6 SFT command is now:
cargo run -p psionic-train -- sft \
--config configs/legal/qwen36_sft_smoke.json
That smoke config uses Qwen/Qwen3.6-27B metadata, QLoRA-style adapter
settings, assistant-only loss flags, and the all-linear Qwen3.6 target-module
declaration. The current executable training surface is intentionally smaller:
it trains an adapter-only LM-head LoRA update over tiny legal hidden-state
samples, writes adapter.safetensors, loss_curve.json,
checkpoint_summary.json, and training_receipt.json, and records that no
Python process or Python-generated trainer artifact was used. It proves the
Rust config, training, receipt, and export loop. Full dense Qwen3.6 causal-LM
target coverage remains the next model-path expansion work.
The first Rust-only Qwen3.6 DPO command is now:
cargo run -p psionic-train -- dpo \
--config configs/legal/qwen36_dpo_smoke.json
That smoke config loads the parent SFT LoRA adapter, bootstraps it from
configs/legal/qwen36_sft_smoke.json when the local target artifact is
missing, loads legal_dpo_v1 prompt/chosen/rejected pairs, renders prompts
through the Qwen3.6 direct-answer template, and runs adapter-only weighted
chosen/rejected updates with beta = 0.25. It writes the same core artifact
family as SFT: adapter.safetensors, loss_curve.json,
checkpoint_summary.json, and training_receipt.json.
The checked smoke DPO dataset lives at
fixtures/legal_benchmark/dpo_smoke/legal-dpo-v1.jsonl. The 2026-05-20 local
command run completed 6 Rust-only steps over 22 pairs, moved the synthetic
preference accuracy from 0.59090906 to 0.95454544, and moved the average
chosen-minus-rejected logprob margin from 0.3191057 to 4.2714095. This is
evidence that the adapter-only DPO path can train toward file-writing
preference behavior on the synthetic smoke surface. It is not a score claim on
private Harvey tasks.
The same DPO adapter path was accepted by the deterministic replay eval suite:
harvey_public_three_deterministic_replay_v1 reported base 3333 bps,
adapter 10000 bps, and delta 6667 bps with report hash
bd01ce5a8653414a2189d935c80c835c774f55f195ed6809021c135a352faa66. That is
replay-harness compatibility evidence, not retained benchmark proof.
The first verifier reward-trace builder for GRPO-style training is:
cargo run -p psionic-eval --example legal_benchmark_build_reward_traces -- \
--runs target/legal/reward_trace_public_three \
--out target/legal/reward_trace_public_three/legal-reward-v1.jsonl
The builder recursively finds run directories with run_record.json, loads
optional score_report.json, run_receipt.json, task_spec.json, and output
manifests, and emits one deterministic JSONL trace per run. It computes
workflow reward from receipts and verifier artifacts only: required file
presence, correct path, non-empty answer, answer length, source-use evidence,
submission state, answer integrity, public score delta, and hidden-leakage
status. It keeps workflow reward separate from the legal-content score so RL
can train on file/tool behavior without confusing that with judge quality.
Hidden or private scoring leakage is a fatal exclusion. Harness-assisted answer content is also a fatal exclusion. Missing required files are not fatal; they stay in the dataset with low reward so GRPO can learn from them.
The 2026-05-20 local public-three replay command:
cargo run -p psionic-eval --example legal_benchmark_eval_suite -- \
--suite suites/harvey_public_three.json \
--model Qwen/Qwen3.6-27B \
--adapter qwen36-legal-reward-smoke \
--out target/legal/reward_trace_public_three
reported base 3333 bps, adapter 10000 bps, delta 6667 bps, and report
hash 3e382da304b9eb756675fc3daf9a68a7c0d41ae560def0ac710d1229d930fa5b.
Building reward traces from that output produced 6 traces, 0 fatal exclusions,
and dataset hash e85352b832dceed154282ff5e6b0de510b6af808b0cce945140387b7710a3e26.
The three adapter-pass traces scored total reward 12; the base missing-answer
and tool-failure traces scored total reward 0. Source-use reward stays 0
on this deterministic replay because these replay records do not include
model-authored document-read tool evidence.
The current honest Harvey MFN local result is run 016: the actual local Qwen
LoRA adapter 005 submitted through the Rust tool loop, wrote its own output,
and scored 4 / 18 on a rubric-free legal work-product proxy. Broad suite
runs 019 and 025 compare model-only and scaffold-assisted prompts across three
public Harvey tasks with runner output mutation disabled. Adapter 020 is a
real local MLX LoRA fine-tune over clean no-cheat supervised trajectories, but
it did not yet improve the broad suite score.
Deterministic Replay Eval
crates/psionic-eval/src/legal_benchmark_eval_suite.rs adds the first local
replay harness for comparing a base model binding and an adapter binding
against the exact same suite. The harness freezes:
- suite id and eval mode
- fixed task order
- source document hashes
- prompt template hash
- scorer version
- inference settings
- base model id and adapter id or adapter artifact hash
The checked smoke suite from the first pass is still available as:
suites/harvey_public_three.jsonfixtures/legal_benchmark/eval_suite_public_three/*
The current public suite catalog is broader and can be listed with:
cargo run -p psionic-eval --example legal_benchmark_list_suites
Materialized suite names:
harvey_public_001_single: one task, smoke only.harvey_public_003_workflow: three frozen workflow tasks.harvey_public_010_mixed: ten public-style tasks across review, analysis, classification, calculation, and drafting workflows.
Reserved suite names:
harvey_public_025_regressionharvey_public_050_candidate_gate
Those two names are intentionally not materialized yet. They are placeholders for future real public tasks, not padded copies of the current ten-task set.
Run it with:
cargo run -p psionic-eval --example legal_benchmark_eval_suite -- \
--suite suites/harvey_public_three.json \
--model Qwen/Qwen3.6-27B \
--adapter target/legal/qwen36_sft_smoke/adapter.safetensors \
--out runs/harvey-public-three-smoke
Named suites can also be run directly:
cargo run -p psionic-eval --example legal_benchmark_eval_suite -- \
--suite harvey_public_003_workflow
cargo run -p psionic-eval --example legal_benchmark_eval_suite -- \
--suite harvey_public_010_mixed
The output directory contains:
eval_report.jsonpromotion_gate_input.jsonreplay_receipt.json- per-task base and adapter run records and score reports
- base and adapter static Markdown summaries
The report separates answer-file write rate, correct answer-path rate, submit rate, integrity-valid rate, mean legal score, median legal score, per-task score, score by task type, failure-class counts, and candidate regression against the champion binding. It also writes a promotion-gate JSON object so later adapter registry work can make a single promote/hold/reject decision without re-parsing scorer internals.
Recorded local smoke results on 2026-05-20:
harvey_public_003_workflow: base3333bps, adapter10000bps, delta6667bps, median adapter10000bps, report hash47b7125199cb642e550a4133938b3dc30de031c64768894848da085e5e1b4636.harvey_public_010_mixed: base3000bps, adapter10000bps, delta7000bps, median adapter10000bps, report hash704198a15af55cf5aa742dcec1861b65baeb577b62b4f35d9c5bd5db6b013df3.
Plain boundary: this harness replays declared local outputs through the Rust scorer. It is useful for proving that the evaluator, receipts, ordering, and promotion inputs are stable. It is not proof that a model improved on hidden Harvey tasks. Hidden audit suites are rejected if marked training-allowed.
Distributed Adapter Merge
crates/psionic-train/src/qwen_legal_lora_merge.rs adds the merge step between
Pylon worker training and benchmark evaluation. It reads a JSON manifest,
verifies each worker adapter artifact hash, loads real LM-head LoRA
safetensors files, and writes one aggregate adapter plus a receipt.
Run the retained smoke merge with:
cargo run -p psionic-train -- merge-lora \
--manifest merge/legal-sft-round-001.json
The receipt records the parent adapter hash, worker adapter hashes, dataset
shard hashes, token counts, merge weights, output adapter hash, validation
metrics from suites/harvey_public_three.json, and the deterministic replay
command. The command supports both token-weighted delta averaging and
sequential shard handoff. The promotion gate is explicit and conservative: the
merged adapter is only promotable if the same-suite local eval beats the
declared champion and has no integrity, tool, or timeout failures.
Adapter Registry And Promotion Gates
crates/psionic-train/src/qwen_legal_adapter_registry.rs adds the first local
Qwen legal adapter registry. Each registry entry records the adapter id, base
model id and hash, training dataset id and hash, training config id and hash,
Psionic version, git commit, training workers, training receipt hash, eval
suite id and hash, eval result hash, parent adapter id, promotion status, and
the eval summary used by hard gates.
Register adapters with:
cargo run -p psionic-train --example qwen_legal_register_adapter -- \
fixtures/legal_benchmark/adapter_registry/qwen_legal_champion_adapter_manifest.json
cargo run -p psionic-train --example qwen_legal_register_adapter -- \
fixtures/legal_benchmark/adapter_registry/qwen_legal_candidate_adapter_manifest.json
Promote a candidate with:
cargo run -p psionic-train --example qwen_legal_promote_adapter -- \
--candidate qwen36-legal-public-three-candidate-001 \
--suite harvey_public_three_deterministic_replay_v1
Set PSIONIC_LEGAL_ADAPTER_REGISTRY or pass --registry <path> to use a
non-default local registry path. The default path is
target/legal/qwen_adapter_registry/registry.json.
Registration rejects missing training receipts, missing eval receipts, empty worker sets, excluded training data, invalid hashes, hidden benchmark leakage, harness-modified answer text, integrity failures, and adapters that were not produced by an allowed Psionic/Pylon path. Promotion rejects lower scores, different suite hashes, answer-file write-rate regressions, required-workflow regressions, hidden leakage, incomplete receipts, and non-Psionic production paths. A promoted candidate supersedes the previous champion for that suite and writes a promotion receipt next to the registry.
Artifact Manifests And Hashing
The module exposes SHA-256 helpers over canonical serde JSON encodings:
task_spec_digestartifact_manifest_digestrun_config_digestrun_record_digesttranscript_digestscore_report_digestcomparison_report_digest
These helpers are namespaced by artifact family. Downstream code should store the digest and the JSON artifact together rather than recomputing identity from partial database rows.
The module also provides reusable manifest helpers:
artifact_from_filebuild_input_artifact_manifestbuild_output_artifact_manifestbuild_derived_artifact_manifestcompare_artifact_manifests
Input manifests are built from normalized task source artifacts. Output and derived manifests take generated artifacts from the runner or extractor layer. All builders sort artifacts deterministically before hashing, so unchanged inputs produce stable manifest hashes and changed source bytes produce manifest drift that can be compared before a run is trusted.
Fixture
The minimal checked fixture lives at:
fixtures/legal_benchmark/minimal_task_bundle.json
It covers one Harvey-compatible legal task, input and output manifests, a run config, a run record, a score report, and a comparison report. The fixture is used by unit tests to prove the schema round-trips and the digest helpers are deterministic.
Harvey Compatibility Loader
The module crates/psionic-eval/src/legal_benchmark_harvey.rs scans a Harvey
tasks directory into owned BenchmarkTaskSpec values and a
HarveyCorpusSummary.
The loader:
- discovers nested
task.jsonfiles under a caller-supplied tasks root - parses Harvey fields
title,work_type,tags,instructions,deliverables, andcriteria - reads sibling
documents/files as source artifacts - preserves the upstream commit and task path under
SourceCompatibility - validates missing documents directories, empty criteria, empty deliverables, and criteria that reference unknown deliverables
- reports task, practice-area, criterion, source-document, deliverable, work-type, and extension counts
Run the summary scanner with:
cargo run -p psionic-eval --example legal_benchmark_harvey_scan -- \
/Users/christopherdavid/work/competition/repos/harvey-labs/tasks \
5aa41694
For the audited Harvey checkout, the expected summary is:
- 1,251 tasks
- 24 practice areas
- 74,990 criteria
- 9,537 source documents
Sandbox Boundary
The local sandbox boundary for extraction and tool execution is documented in
docs/PODMAN_SANDBOX_BACKEND.md and implemented in
crates/psionic-sandbox/src/podman.rs.
For legal benchmark runs, the default Podman config disables network access,
mounts source documents read-only at /workspace/inputs, and exposes writable
scratch and output paths at /workspace/scratch and /workspace/output.
Path validation canonicalizes host roots and rejects traversal or symlink
escapes before a container command is built.
Document Extraction
The extraction contract lives in
crates/psionic-eval/src/legal_benchmark_extraction.rs.
It defines:
ArtifactExtractorArtifactExtractionPolicyArtifactExtractionResultExtractionReceiptExtractionCoverageExtractionFailureKindArtifactExtractorRegistry
Native extraction handles text, Markdown, JSON, CSV/TSV, XML/HTML, YAML, TOML, log-like UTF-8 inputs, DOCX, PPTX, XLSX, and EML without external tools. The Office path reads the Open XML ZIP container and preserves text from supported document, slide, workbook, shared-string, header, and footer XML parts. The EML path preserves common headers and readable body text. Both native paths are lossy by design and emit extraction warnings because layout, attachments, tracked changes, formulas, comments, MIME encodings, and hidden workbook semantics still require a higher-fidelity sandboxed extractor.
The registry still declares pinned sandboxed external adapter specs for PDF and
future high-fidelity Office/EML extraction. Until a live sandbox command
executor is attached, those external adapters return structured
external_tool_unavailable or policy-denied receipts instead of panicking or
crashing a sweep.
The 2026-05-20 local run against the audited Harvey checkout proves the native path across the full corpus:
cargo run -q -p psionic-eval --no-default-features \
--example legal_benchmark_extract_slice -- \
/Users/christopherdavid/work/competition/repos/harvey-labs/tasks \
5aa41694 1251
Result: 1,251 tasks scanned, 9,537 source artifacts extracted, and zero
structured extraction failures. The prior 25-task pre-change slice extracted
1 of 198 artifacts and returned 197 ExternalToolUnavailable failures.
Run records and score reports now retain extraction_receipt_refs so operator
surfaces can distinguish extraction failure, missing content, and bad
reasoning.
Tool Surface
The closed Rust tool set is documented in docs/LEGAL_BENCHMARK_TOOLS.md and
implemented in crates/psionic-eval/src/legal_benchmark_tools.rs.
It covers shell, read, write, edit, glob, and grep with typed inputs, typed outputs, structured errors, byte metrics, touched paths, transcript events, and tool-call records. Read can prefer extracted text from the extraction layer; write and edit are restricted to workspace/output roots; shell remains sandbox-owned and routes through the Podman backend only when the full-feature caller attaches one.
High-score document helpers now extend the same receipt-backed surface with inventory, EML summary, spreadsheet summary, page-targeted PDF search, evidence table construction, and deliverable validation. These tools improve source coverage, evidence traceability, and pre-judge output checks without exposing hidden criteria. Normalized Harvey tasks enable this full document-tool set by default so live runs do not fall back to the older read/write/grep-only surface.
Provider Adapter Layer
The provider-neutral model contract is documented in
docs/LEGAL_BENCHMARK_PROVIDERS.md and implemented in
crates/psionic-eval/src/legal_benchmark_provider.rs.
It defines model requests, messages, tool specs, tool calls, tool-result
messages, usage accounting, structured provider failures, retry policy, a
Google Vertex Gemini generateContent adapter for the active
gemini-3-flash-preview benchmark lane, OpenAI-compatible and Anthropic
protocol adapters for fallback/parity lanes, and deterministic CI mocks. Routes
record provider family, model id, model config hash, elapsed time, retry count,
raw response hash, and secret reference id without writing raw credentials into
run artifacts.
Agent Runner
The Rust agent loop is documented in docs/LEGAL_BENCHMARK_RUNNER.md and
implemented in crates/psionic-eval/src/legal_benchmark_agent.rs.
It builds policy/task prompts, drives provider turns, executes tool calls,
requires explicit JSON submit/finalize semantics, classifies terminal states,
and writes config.json, transcript.jsonl, metrics.json,
output_artifact_manifest.json, extraction_receipts.json,
tool_receipts.json, run_record.json, and run_receipt.json.
Integrity mode is the default prompt policy. The runner does not render hidden criteria into model-visible messages; hill-climb runs may render only policy-approved derived checklist items from run config metadata.
Coverage Tracker
Criterion-adjacent coverage tracking is documented in
docs/LEGAL_BENCHMARK_COVERAGE.md and implemented in
crates/psionic-eval/src/legal_benchmark_coverage.rs.
Run records persist CoverageSnapshot values covering discovered/read
documents, extracted facts, evidence refs, drafted deliverable sections,
validations, self-checks, and the policy mode used. Score reports copy the
snapshot and add post-judge failure comparisons that classify missed criteria
as coverage, extraction, drafting, or reasoning gaps for Autopilot4 import.
Evaluator And Judge Interface
The criterion-scoped evaluator is documented in
docs/LEGAL_BENCHMARK_EVALUATOR.md and implemented in
crates/psionic-eval/src/legal_benchmark_evaluator.rs.
It loads completed task/run/output-manifest artifacts, runs deterministic
manifest and deliverable prechecks, extracts output text per criterion, calls a
provider-neutral judge adapter, and emits ScoreReport values with all-pass
status, criterion pass rate, judge provenance, confidence, latency, cost,
document coverage, failure diagnostics, extraction receipt refs, coverage
snapshots, and missed-criterion failure classifications.
Static Reports
Static reporting is documented in docs/LEGAL_BENCHMARK_REPORTS.md and
implemented in crates/psionic-eval/src/legal_benchmark_reports.rs.
It generates Markdown reports for humans plus stable Autopilot4 import JSON
with global, per-task, per-model-config, comparison, and failure-cluster
summaries. The example command
cargo run -p psionic-eval --example legal_benchmark_report -- <score.json> <out-dir>
writes report.md, autopilot_report.json, and failure_clusters.json
without live provider credentials.
Sweep Runner
Sweep planning and manifests are documented in docs/LEGAL_BENCHMARK_SWEEPS.md
and implemented in crates/psionic-eval/src/legal_benchmark_sweeps.rs.
The sweep layer plans task/config jobs, expands provider/reasoning/context/ extraction/tool-policy matrices, applies resume state, enforces cost, wall-time, token, and failure budgets, keeps going through individual task/model failures, and emits a manifest with skipped, resumed, succeeded, failed, blocked, and budget-exhausted job states for Autopilot4 import.
Matrix exports summarize every recorded config hash by all-pass score, criterion pass rate, document coverage, reliability, cost, and latency, then mark Pareto-front configs for promotion-gate review.
Product Regression Guardrails
Product regression guardrails are documented in
docs/LEGAL_BENCHMARK_REGRESSION_GUARDRAILS.md and implemented in
crates/psionic-eval/src/legal_benchmark_regression.rs.
The guardrail suite uses synthetic fixtures for chat, Coder, Work Orders, GitHub provider, CRM, memory, and provider/tool routing. Gate reports include both benchmark target scores and product regression scores, export Autopilot4 release-gate import JSON, create blocking Work Orders for failed product regressions, and disallow live user data or Harvey hidden criteria in the regression fixture suite.
CI And Golden Fixtures
Repo-native compatibility checks are documented in
docs/LEGAL_BENCHMARK_CI.md and run with
scripts/check-legal-benchmark-ci.sh.
The check target pins the audited Harvey corpus metadata, verifies the minimal normalization snapshot, covers sandbox traversal and symlink escape behavior, and exercises mock report, sweep, and product-regression fixtures without live provider credentials.
Next Work
The next implementation issue is the Autopilot4-side release-gate import and operator surface for these Psionic reports.