Eval reports

September 20, 2026 · View on GitHub

Read this when…ReportStatus
You want the current headline: Spark vs AMD Strix Halo (Vulkan, ROCm) vs hosted Terra on gpt-oss-120beval-report-2026-09-spark-halo-terra.mdlatest (Sept 13, 2026)
You need Spark/Terra detail: per-tag and per-company tables, the 20b model, every miss itemised, costeval-report-2026-09-spark-cuda-rerun.mdcurrent for the Spark and Terra columns
You want Qwen3.8-27B on the same Spark suite, including why it was ~5× slower than gpt-oss-120beval-report-2026-09-spark-cuda-qwen.mdQwen add-on (Sept 15, 2026)
You want Nemotron-3-Super on the same suite (131K-budget vs full 1M window)eval-report-2026-09-spark-cuda-nemotron.mdNemotron add-on (Sept 16–17, 2026)
You want Qwen3.8-27B on a single RTX 5090 (Windows, CUDA): the 131K replication and the native-262K run that recovers Goldman Sachseval-report-2026-09-rtx5090-cuda-qwen.mdQwen on RTX 5090 (Sept 15, 2026)
You want the first FinanceBench and coding tier-1 (HumanEval/MBPP) runs, Qwen3.8-27B on the RTX 5090, reasoning on vs off, and the harness fixes they neededeval-report-2026-09-rtx5090-financebench-swe-qwen.mdfirst runs of both suites (Sept 15, 2026)
You want PrismML's Ternary Bonsai 2 27B (Qwen3.8-27B at 1.7 bits/weight, 7 GB) against the 4-bit Qwen on the same RTX 5090: extraction at 131K and 262K, FinanceBench, coding tier 1eval-report-2026-09-rtx5090-bonsai2.mdternary add-on (Sept 18, 2026)
You want Gemma 4 31B (thinking on vs off) and Poolside's Laguna S 2.1 on the Spark: why fatter tokenizers push Goldman into chunks, and the fastest local run after gpt-oss-120beval-report-2026-09-spark-cuda-gemma-laguna.mdGemma / Laguna add-on (Sept 19–20, 2026)
You want the pre-fix run and the explanation of what the harness fixes changedeval-report-2026-09-spark-cuda.mdsuperseded; kept for the cross-check

Each number has one home: the Halo report owns the AMD columns, the re-run owns the Spark and Terra columns, and the first report owns the pre-fix baseline. The machine-generated table over all committed runs is state/evals/report-sec.md (uv run local-llm eval report --suite sec regenerates it).

inference-stack-spark-vs-halo.png is a binary export; there is no diagram source in the repo.

Every report now carries a "Review notes" block under its Caveats listing what the numbers cannot support (single run, no memorisation control yet, unequal context, revised ground truth).

Per-model reference cards (identity, exact serving config, measured results, strengths and weaknesses, changelog) live in models/. Update a card whenever a run lands.

A talk-through slide deck of the results lives in deck/ (local-inference-2026-09.pdf; Marp source alongside).

Posts and write-ups published outside the repo, with links and the feedback they drew, are logged in posts/README.md.