Harness Engineering Study Guide

August 27, 2026 · View on GitHub

中文 | English

License: MIT Articles Translations Read online

Harness Engineering Study Guide

A deep-dive learning archive on Harness Engineering — from concept to practice

Harness Engineering — humans steer, agents execute (intro deck cover, 2026-08 snapshot)

Cover from the repo's own intro deck + poster — generated with the open-kimi-ppt skill; editable PPTD sources live in this repo

Introduction

This is an evolving learning project. Harness Engineering is an engineering paradigm proposed by OpenAI in February 2026: engineers stop writing code and instead design environments, clarify intent, and build feedback loops so AI agents can work reliably.

Humans steer. Agents execute.

This repository documents the full learning journey — from reading the original article, breaking down concepts, forming independent thoughts, hands-on experiments, to producing shareable work. We hope it helps others exploring AI-native engineering.

Source: OpenAI — Harness Engineering: Harnessing Codex in an Agent-First World

Note: The insights shared here are not universally applicable. Please adapt them to your own context.

⚡ In One Sentence

Traditional:          Humans write code → Machines run code
Harness Engineering:  Humans design constraints → Agents write code → Machines run code

The core shift: an engineer's output moves from code to constraint systems — AGENTS.md, architecture rules, custom linters, and feedback loops.

🧭 Six Core Concepts

1. Repo as System of Record — If it's not in the repo, it doesn't exist for the agent

Slack threads, Google Docs, knowledge in people's heads = invisible to the agent. All decisions, specs, and plans must be committed as versioned artifacts.

→ See concepts/01-repo-as-source-of-truth.md

2. Map, Not Manual — AGENTS.md is a table of contents, not an encyclopedia

A ~100-line entry file pointing to deeper docs. Progressive disclosure: the agent starts from a small, stable entry point and is guided where to look next. Three ways a giant instruction file fails: crowds out context, impossible to maintain, can't be mechanically verified.

→ See concepts/00-overview.md

3. Mechanical Enforcement — Docs rot; lint rules don't

Custom linters + structural tests = invariant guardians. Lint error messages embed fix instructions so agents can self-correct. Enforce boundaries centrally, allow autonomy locally.

→ See concepts/02-mechanical-enforcement.md

4. Agent Readability — Optimize for the agent's ability to reason

Prefer "boring" technologies (stable APIs, well-represented in training data). Sometimes re-implementing a focused subset is cheaper than wrapping opaque upstream behavior. Make the app launchable per git worktree.

→ See concepts/04-agent-readability.md

5. Throughput Changes Merge Philosophy — Correction is cheap; waiting is expensive

Short PR lifecycles. Flaky tests resolved by re-runs rather than blocking indefinitely. In a system where agent throughput far exceeds human attention, this is usually the right call.

→ See concepts/05-throughput-changes-merge.md

6. Entropy Management = Garbage Collection — Tech debt is a high-interest loan

Agents reproduce existing patterns in the repo — including bad ones. Codify "golden rules" into the repo. Run periodic background tasks to scan for drift, update quality scores, and open targeted refactoring PRs.

→ See concepts/03-entropy-and-garbage-collection.md

🔑 Key Data Points

MetricData
Team size3 → 7 engineers
Time span5 months
Codebase~1 million lines
PRs merged~1,500
PRs per engineer per day3.5 (still growing after scaling)
Single run duration6+ hours (often during human sleep)
Efficiency estimate~1/10 of manual coding time

📂 Repository Structure

harness-engineering/
├── README.md              ← Chinese (primary)
├── README.en.md           ← You are here
├── AGENTS.md              ← Repo navigation entry (for agents)

├── concepts/              # Phase 1: Concept notes (8 articles)
│   ├── 00-overview.md     #   Overview of all six concepts
│   ├── 01-repo-as-...     #   Repo as source of truth
│   ├── 02-mechanical-...  #   Mechanical enforcement
│   ├── 03-entropy-...     #   Entropy & garbage collection
│   ├── 04-agent-...       #   Agent readability
│   ├── 05-throughput-...  #   Throughput changes merge philosophy
│   ├── 06-harness-...     #   Harness definition (Fowler control-theory extension)
│   └── 07-spec-as-product.md #   Spec as product (Symphony extension)

├── thinking/              # Phase 2: Independent analysis (11 articles)
├── practice/              # Phase 3: Hands-on experiments (1 Ralph Demo)
├── feedback/              # Phase 4: Lessons learned (1 article)
├── works/                 # Phase 5: Shareable outputs (40 translations + 1 original + 2 external Chinese captures)
├── tools/                 # Tools that reduce the 6 complexity dimensions
├── prompts/               # Validated prompts collection
└── references/            # External resource index (79 articles with deep summaries)

Each subdirectory has its own AGENTS.md explaining its purpose and conventions — a direct practice of the "progressive disclosure" principle from the original article.

🚀 Learning Path

  • Phase 1: Understand core concepts — 8 concept notes covering OpenAI's six concepts + Fowler's control-theory extension + Symphony's spec-as-product
  • Phase 2: Form your own opinions — 11 independent analyses (ongoing)
  • Phase 3: Pick a small project to practice — Ralph Demo completed (321s, $0.31)
  • Phase 4: Record feedback & iterations — 1 article (ongoing)
  • Phase 5: Produce shareable work — 40 professional translations + 1 original synthesis + 2 external Chinese captures

📚 Research Library

79 articles across three knowledge tracks + 2 extended readings:

TrackCoveragePerspectives
AI-Era Harness Engineering75 articlesOpenAI → Fowler → Anthropic → LangChain → Stanford → Claude Code reverse engineering & source leak → Subagent runtime → Sensors/SPDD/ADLC → Out-of-scope, safety auditing & quality postmortems → Evaluation trilogy → Dynamic workflows → Origins (Ralph / Hashimoto) & discipline synthesis → Codex harness anatomy → Loop Engineering trilogy → Self-evolving harnesses & RSI → Formal verification → Multi-agent scaling (Cursor / C compiler) → Official containment & evals methodology → Behavior maps / DSLs / local models / outer-loop accountability → industrial-scale mechanical porting (Bun) & harness-model co-evolution (HarnessX) → long-running harness foundations & eval-environment confounders (Anthropic backfill) → harness operations metrics & reward hacking (Cursor backfill) → tool schemas are not neutral → the software-factory debate (Dex Horthy / Osmani) → agent-swarm cost economics → deleting 80% of the system prompt → a code-review-sensor benchmark (ReviewBench) → empirical refutation of TDD-as-process (Böckeler) → practical loop engineering (Osmani) → an org-scale adoption snapshot (Zalando) → open neutral harnesses & white-box compaction (Pi duo) → frozen-artifact cross-model transfer (StarHarness)
Cloud-Native Harness.io2 articlesCI/CD platform architecture (same name, different meaning)
Efficiency Paradox & Capability Evolution2 articlesYDD systematic teardown + METR follow-up (measurement-methodology crisis)
Extended Reading2 articlesContext Engineering, Human-Agent collaboration

See references/articles.md — each article includes core thesis, key data, and cross-article connections.

📖 Translations

40 Chinese translations of key articles (click to expand)
TranslationOriginal AuthorSource
Eight Years of WantingLalit MagantiPersonal blog
Evaluating code review agents with ReviewBenchNick HollonLangChain
TDD inside the agent loop - theater or actual value?Birgitta Böckelermartinfowler.com
Practical Loop EngineeringAddy OsmaniAddyOsmani.com
Agentic Engineering at Zalando: A SnapshotBartosz OcytkoZalando Engineering
What Is a Harness?Earendil / Pi teamearendil.com
How Compaction Works in PiEarendil / Pi teamearendil.com
StarHarness: Evolving Harnesses with Stratified SearchServiceNow / Mila et al.arXiv
The New Rules of Context Engineering for Claude 5Thariq ShihiparAnthropic / Claude
Better Models: Worse ToolsArmin RonacherPersonal blog
Rewriting Bun in RustJarred SumnerBun Blog
Building a C Compiler with a Team of Parallel ClaudesNicholas CarliniAnthropic
Scaling Long-Running Autonomous CodingWilson LinCursor
How We Contain Claude Across ProductsMax McGuinness et al.Anthropic
Harness Engineering for Self-ImprovementLilian WengLil'Log
Loop EngineeringAddy OsmaniPersonal blog
The Coming LoopArmin RonacherPersonal blog
A Harness for Every Task: Dynamic WorkflowsThariq Shihipar et al.Anthropic / Claude
METR: Changing Our Productivity Experiment DesignJoel Becker et al.METR
Inside the ScaffoldBenjamin RombautHuawei / arXiv
Meta-HarnessYoonho Lee et al.Stanford / arXiv
Harness Engineering (full)Birgitta BöckelerMartin Fowler
Harness Engineering (memo)Birgitta BöckelerMartin Fowler
Encoding Team StandardsRahul GargMartin Fowler
Feedback FlywheelRahul GargMartin Fowler
Scaling Managed AgentsLance Martin et al.Anthropic
Agent Evaluation ChecklistLangChain TeamLangChain
Agent-driven DevelopmentTyler McGoffinGitHub
Continual LearningHarrison ChaseLangChain
Codex Orchestration Spec: SymphonyKotliarskyi et al.OpenAI
Claude Code Architecture (Reverse Engineered)Vikash RungtaSubstack
Maintainability Sensors for Coding AgentsBirgitta BöckelerMartin Fowler
Structured-Prompt-Driven Development (SPDD)Wei Zhang et al.Martin Fowler
The Agent Development Lifecycle (ADLC)Harrison ChaseLangChain
Interpreters in Deep AgentsHunter LovellLangChain
Claude Code Quality PostmortemAnthropic EngAnthropic
Agentic Harness Engineering (paper)Jiahang Lin et al.Fudan / arXiv
Overeager Coding Agents (paper)Yubin Qu et al.arXiv
How I Use AI to CodeChris ParsonsPersonal blog
How We Built LangSmith EnginePalash ShahLangChain

Source Material

ResourceDescription
OpenAI Original ArticleThe full Harness Engineering exposition

Ralph Series — Harness Engineering in Practice

The "Ralph Wiggum Loop" is the core implementation pattern of Harness Engineering: agents work autonomously in a loop until the task is complete.

ProjectStarsDescription
snarktank/ralph13.6kOriginal Ralph: bash script that repeatedly spawns AI with fresh context until all PRD items pass. 6 core tenets
ralph-orchestrator2.3kRust evolution: Hat-based personas + event-driven coordination + multi-backend (Claude/Kiro/Gemini/Codex) + backpressure gates + persistent memory
bmad-ralph2BMAD method + Ralph: parallel Claude Code worktrees + three-layer self-healing (retry → restart → diagnose) + SQLite state machine

Ralph Tenets ↔ Harness Engineering Mapping

Ralph TenetHarness Engineering Concept
Fresh Context Is ReliabilityAgent Readability — re-read everything each iteration
Backpressure Over PrescriptionMechanical Enforcement — don't prescribe how; gate bad output
The Plan Is DisposableEntropy Management — regeneration costs one planning loop
Disk Is State, Git Is MemoryRepo as System of Record — files are the handoff mechanism
Steer With Signals, Not ScriptsHumans Steer — add signs, not scripts
Let Ralph RalphAgents Execute — sit on the loop, not in it

Community & Extended

ResourceDescription
vibe-coding-cnChinese Vibe Coding community guide
Mitchell Hashimoto: Engineer the HarnessWhere "harness engineering" got its name (now indexed as article #29 in references/articles.md)

🛠️ Development Notes

The repo ships with a consistency checker, scripts/check-consistency.sh, guarding against count and fidelity drift across fourteen layers of checks:

  • C1-C2references/articles.md$ \text{article} \text{count} + \text{its} 4 \text{downstream} \text{claim} \text{sites} (\text{README} \times 2 \text{badges}, $prompts/deep-research-tracker.md header, references/AGENTS.md overview)
  • C3 — actual *.md file counts in concepts/ / thinking/ / feedback/ match the README "X 篇" claims
  • C4works/*-translation.md file count matches every translation-count claim (badges, table summaries, Phase 5 mentions, AGENTS snapshot, table row counts)
  • C5 — the README structure tree lists every single concepts/*.md file
  • C6 — the "不计入 N 篇" exclusion note at the end of `references/articles.md$ \text{matches} \text{the} \text{C1} \text{authority} \text{count}
  • \text{C7} — \text{per}-\text{track} \text{counts} (\text{Track} 1/2/3) \text{stay} \text{consistent} \text{across} \text{their} 4 \text{downstream} \text{claim} \text{sites} (\text{README} \text{research}-\text{library} \text{tables} \times 2, $references/AGENTS.mdtrack headings,prompts/deep-research-tracker.md` track lines)
  • C8 — local translation-pipeline guard: once translate/<...>/sources/<slug>/source-full.md is captured, the matching 01-analysis.md may no longer claim "abstract-only / fetch full text later". translate/ is gitignored, so this auto-SKIPs on CI and clean clones
  • C9 — authored prose in concepts/ / thinking/ / feedback/ must not restate library counts ("N articles / N translations") as live facts; historical mentions must carry a dated-snapshot qualifier, otherwise drop the number and link references/articles.md
  • C10 — figure fidelity (purely local, zero network): every translation's frontmatter must declare sourceFigureCount, the body must embed at least that many images, and every local embed path must exist on disk (null = source unavailable / unaudited → SKIP)
  • C11 — markdown table shape: in the checked files, every table row must carry the same cell count as its header
  • C12 — every numbered entry in references/articles.md must carry the 作者: and 日期: fields
  • C13 — zero-figure claims need an audit trail. C10 can only falsify OVER-claiming, so sourceFigureCount: 0 is unfalsifiable locally — that hole shipped a false 0 on 2026-07-27 (the source had 4 body figures). Any translation claiming 0 must therefore also carry sourceFigureAudit containing a YYYY-MM-DD date, stating how the claim was verified
  • C14 — docs-site harness integrity: the VitePress sidebar and every displayed count must be derived from the filesystem at build time by .vitepress/sidebar.mjs; site sources (index.md, .vitepress/**) must not hardcode counts, and node .vitepress/sidebar.mjs --verify asserts every first-class content file appears in the generated sidebar exactly once and rejects any symlink in the repo (symlinks can leak external files into the published artifact). On the artifact side, scripts/verify-dist.mjs asserts published pages and their .md copies correspond one-to-one, that relative links and images inside the copies resolve, and that dist contains no symlinks

Enable the pre-commit hook after first clone:

git config core.hooksPath .githooks

Once enabled, every commit touching the README, AGENTS.md, references/articles.md, references/AGENTS.md, index.md, .vitepress/, scripts/check-consistency.sh, or any *.md (nested included) under concepts/ / thinking/ / feedback/ / works/ / practice/ / tools/ / prompts/ runs the checks automatically; staging a symlink at any path is rejected outright (the repo-wide C14 ban, judged on the staged state). Unrelated commits are left alone.

Run manually: bash scripts/check-consistency.sh

CI backstop: even without the local hook, GitHub Actions (.github/workflows/consistency.yml) runs the same script on every push / PR (no path filters, so the branch-protection required check always gets reported). The local hook is fast feedback during development; CI is the actual merge gate.

See the "机械化检查" section of the root AGENTS.md for details.

🪞 The repo is its own harness (self-reference)

This archive now curates itself.

Bringing in outside research no longer runs on vibes — it follows a pipeline frozen into a skill, curate-research: review is automated by parallel agents (the feedback loop), scripts/check-consistency.sh keeps counts and fidelity from drifting via C1–C14 (the mechanical rail), and whether something gets in is always a human gate (humans steer, agents execute).

So the constraints themselves became the product — exactly what concepts/07-spec-as-product.md argues, except this time the subject is the repo itself.

🤝 Contributing

Contributions via Issues and PRs are welcome:

  • Add concept notes (concepts/ has gaps to fill)
  • Share your independent thinking (thinking/)
  • Contribute practice cases (practice/)
  • Recommend related resources (references/)

📞 Contact

ChannelLink
GitHub@deusyu
X (Twitter)@0xdeusyu
Telegram@DeusThink
Telegram Group@talkdeusyu
Telegram Channel@lovedesuyu
Emailrainman.deus@gmail.com

Star History

If you find this project helpful, please consider giving it a Star ⭐!

Star History Chart

If this learning archive has saved you time, consider sponsoring my open-source work — your support keeps it updated, free, and open.

Sponsor

📄 License

MIT