Contributing to buildaharness
September 18, 2026 · View on GitHub
Current state
v0.8.0 — fully implemented. All four adapter runtimes are executable. Full 11-layer harness architecture is in place. No open RFCs.
What's shipped:
- FlowSpec schema v1.0.0, 27 canvas node types (14 base + 13 harness), 5 reference flows, ADR-001 closed
- XYFlow canvas, LangGraph + CrewAI + Mastra + MAF adapters, auth, Langfuse, execution, HITL
- Observability stack (ClickHouse + Redis + Langfuse), OTel traces, token counts
- Team RBAC, JWT revocation, offline/online eval, prompt versioning, A2A, deploy, marketplace
- SSO/OIDC + SCIM, Helm chart, Yjs real-time collab,
@buildaharness/canvaspackage - Full harness architecture: 11-layer reasoning and control system, harness tests (P0–P11, P-PC, integration, E2E, invariants)
- Aielia — a personal assistant running the full harness client-side every turn (
@buildaharness/aielia), with a CLI, a browser build (@buildaharness/chat-ui), and a native desktop app (@buildaharness/desktop) - npm packages:
@buildaharness/harness,@buildaharness/runtime,@buildaharness/react,@buildaharness/canvas,@buildaharness/aielia,@buildaharness/proxy
Not in this repo: a few pieces referenced in internal docs are maintained in a private overlay
and aren't part of this public clone — most visibly the coaching-agent example flow and its
persona/prompt fixtures, and a deeper pseudo-code/state-model architecture write-up (the public
docs/architecture.md covers the same system at a level intended for external
contributors). The adapter/agents/coaching/ tool implementations that ship here are public and
usable by any flow via fn_ref; the example flow and session-runner tooling built on top of them
are not.
Architecture Decision Records
Significant design decisions are recorded as ADRs and referenced throughout the code and docs by their ADR-NNN identifier. The records themselves are maintained outside this repo; the parts that are a contract for contributors are reproduced here (see ADR-001 field contracts under Adding an adapter below) and in spec/CHANGELOG.md. Open a new [adr] issue with new evidence to challenge an accepted decision — don't re-litigate in the original thread.
| ADR | Title | Status |
|---|---|---|
| ADR-001 | Codegen semantics: output_key, *_expr, context_from, memory_write.tier | Accepted |
What we need
Adapter runtime feedback — if you use LangGraph, CrewAI, Mastra, or MAF seriously, tell us where the adapter generates wrong or unidiomatic code. Open an [adapter] issue with the spec JSON, generated output, and the correct output.
New node types — what flow pattern can't you express with the 14 existing types? Open a [node-type] issue with a concrete use case, proposed schema shape, and per-adapter mapping for all four runtimes.
Community components — publish tool wrappers to the marketplace via POST /marketplace. Especially needed: database connectors, cloud API wrappers, data transformation tools. See Publishing a marketplace component.
Eval integration feedback — the eval harness (DeepEval + Ragas) and Langfuse LLM-as-judge are shipped. Tell us what's missing: in-flow quality gates, specific metric implementations, CI regression thresholds.
@buildaharness/canvas consumers — if you embed the canvas package in your own tool, open a [canvas-pkg] issue for anything that doesn't fit the current props API.
Good first issues
Scoped, real gaps with a clear finish line.
1 · chat-ui write-approval UI [chat-ui]
write_file / run_shell_command stage a .pending-actions/<id>.json record and
return needs_approval with pendingActionKind. The CLI and desktop resolve it;
packages/chat-ui's ApprovalCard shows the request but not what will be
written — so file tools are CLI/desktop-only for now (package README, "File
access via tools"). Add a diff/content preview to ApprovalCard for
pendingActionKind: 'write' (and the command + resolved cwd for 'shell'),
sourced from the staged record, so a browser user can approve a staged write with
the same information the CLI prints. App.tsx already threads pendingActionId
through handleApprove / handleDeny.
2 · Make cross-run learning legible and deletable [assistant]
The assistant already learns across runs via ExperienceStore
(packages/aielia/src/), but there's no user-facing surface for it —
you can't see what it has generalised or forget a specific lesson. Add a
/experience CLI command (list + delete-by-id) and a short "What Aielia has
learned" section to the package README, sourced from the existing store API. Done
when a user can enumerate stored experience entries and delete one, with tests.
3 · Inverted index for transcript search [assistant]
Transcript search is a linear scan over every stored message
(packages/aielia/src/). Fine for now, O(n) per query as history
grows. Add an in-memory inverted index (token → message ids) built lazily on first
search and updated on append, falling back to the linear scan when absent. Done
when search returns identical results with the index enabled and a benchmark shows
the win past ~10k messages.
To pick one up, open an issue with the matching label describing your approach before starting.
Issue labels
| Label | Use for |
|---|---|
[spec] | Schema changes — new fields, changed types, new constraints |
[node-type] | New node type proposals or changes to existing types |
[adapter] | Questions, bugs, or improvements specific to one runtime |
[adr] | Proposing a new Architecture Decision Record |
[breaking] | Anything that invalidates currently valid flows |
[docs] | README, CHANGELOG, ADR, or in-schema describe() improvements |
[marketplace] | Component publishing, install behaviour, seeder |
[deploy] | REST/MCP/A2A deploy pipeline, shareable URLs, invoke endpoint |
[eval] | Eval harness, LLM-as-judge, Langfuse scoring |
[observability] | Tracing, token counts, Langfuse wiring |
[collab] | Yjs real-time collaboration, presence, offline persistence |
[canvas-pkg] | @buildaharness/canvas npm package — props API, embedding, theming |
[assistant] | Aielia / @buildaharness/aielia — harness bridge, risk gate, tools, memory, CLI |
[chat-ui] | @buildaharness/chat-ui browser build — approval UI, settings, /try |
Making schema changes
spec/schema.ts is the canonical source of truth. spec/schema.json is derived — never edit it directly.
Every schema PR must include:
spec/schema.ts— the change with adescribe()string explaining field semanticssrc/spec/schema.ts— kept in sync (omit.refine()calls on discriminated union members)packages/canvas/src/spec/schema.ts— same sync, same rulepackages/runtime/src/spec/schema.ts— same sync, same rule (runtime's own copy — see that file's header comment for why it doesn't just depend on@buildaharness/canvasor@buildaharness/flow-spec)spec/schema.json— regeneratedspec/CHANGELOG.md— one entry under the appropriate version header- At least one example flow in
flows/demonstrating the change
Run node scripts/check-schema-sync.mjs to verify all copies are aligned.
Spec version follows semver. Minor versions are additive only (new optional fields). Removing or renaming fields requires a major bump and a migration note in CHANGELOG.md.
Writing an adapter
An adapter is a Python module with one public function:
def compile_<runtime>(spec: dict) -> tuple[str, list[str]]:
"""
Returns (code: str, warnings: list[str]).
Warnings are shown in the canvas compile panel.
"""
Register it in adapter/main.py (SUPPORTED_RUNTIMES + compile dispatch) and in adapter/run_api.py if execution is also supported.
If execution requires a sidecar (like Mastra's Node.js runner), add it to docker-compose.yml and document the protocol.
ADR-001 field contracts — all adapters must respect these
| Field | Contract |
|---|---|
output_key | Node returns {output_key: result}. If absent on llm_call, return {} and log a warning. |
query_expr / key_expr / value_expr | Bare JSONPath ($.state.field). Implement _resolve(expr, state) — see langgraph_adapter.py. |
context_from on edges | Map to the runtime's native context mechanism. Generate a descriptive comment block if no native equivalent exists. |
memory_write.tier | Map to the runtime's memory tier API. Generate a comment if no native tier system exists. |
NODE_SUPPORT_MATRIX — mark compat for all four runtimes
Every node type in spec/schema.ts carries runtime compatibility flags: [LG] LangGraph, [CR] CrewAI, [MA] Mastra, [MS] MS Agent Framework. Mark full / partial / missing for each. The canvas shows a warning badge when the selected runtime has partial support.
Testing your adapter
All 5 reference flows must produce syntactically valid output:
import json, pytest
from <runtime>_adapter import compile_<runtime>
@pytest.mark.parametrize("flow_path", [
"flows/01-rag-agent-flow.json",
"flows/02-content-moderation-hitl-flow.json",
"flows/03-parallel-risk-assessment-flow.json",
"flows/04-research-crew-flow.json",
"flows/05-debate-agent-a2a-flow.json",
])
def test_all_flows_compile(flow_path):
spec = json.loads(open(flow_path).read())
code, warnings = compile_<runtime>(spec)
assert code
compile(code, "<test>", "exec") # Python; use tsc for TypeScript adapters
Adding a node type
Touches nine places:
spec/schema.ts— Zod schema entry withdescribe()stringssrc/spec/schema.ts— canvas copy (omit.refine()calls)packages/canvas/src/spec/schema.ts— canvas package copy (same rule)packages/runtime/src/spec/schema.ts— runtime's own copy (same rule); also add/skip an executor inpackages/runtime/src/executors/depending on whether runtime should execute the new type or treat it as a passthrough stubsrc/store/index.ts—NODE_DEFAULTSentrysrc/canvas/nodes/NodeComponents.tsx— the React component, exported by namesrc/canvas/nodes/index.ts— add to thenodeTypesexport mapsrc/components/ConfigPanel.tsx— panel function + entry inPANEL_MAPsrc/canvas/nodes/BaseNode.tsx— icon (NODE_ICONS) and colour (NODE_HEX)
Adding a harness node type
Harness nodes (those inside src/canvas/nodes/harness/) need two additional steps:
adapter/harness/node_compilers.py— add acompile_<type>_node(node)function and register it inHARNESS_NODE_COMPILERSadapter/harness/__init__.py— export the new compile function and updateHARNESS_NODE_COMPILERS
The harness node is then available to all four adapters automatically via gen_harness_preamble().
Canvas contributions
npm install
npm run dev # → http://localhost:3000
npm test # Vitest — must stay green
# Canvas package (packages/canvas/)
npm run build:canvas # lib build → packages/canvas/dist/
npm run test:canvas # canvas-package tests
npm run typecheck:canvas # TypeScript check
The canvas must never break spec round-trip. npm test validates all 5 reference flows through import → canvas state → export → Zod parse.
@buildaharness/canvas package contributions
The canvas package in packages/canvas/ uses a per-instance Zustand store via createStore() — not a module-level singleton. This is intentional so the component is safe to mount multiple times on a page. Any contribution that introduces module-level mutable state will be rejected.
Props changes require updating packages/canvas/src/BuildAHarnessCanvas.tsx, packages/canvas/README.md, and this file's props table if it changes the public API.
Collab contributions
Real-time collab lives in src/collab/. The Yjs document structure (doc.ts) and the bidirectional sync with Zustand (syncToYjs.ts / syncFromYjs.ts) are the core of the layer. Any change that causes the Yjs doc and the Zustand store to diverge will cause split-brain for collaborators — test this carefully.
The y-websocket server is stateless with respect to the flow spec (it only relays CRDT ops). Do not add server-side state to the collab infrastructure.
Working on Aielia & the npm packages
No Docker, no stack. @buildaharness/aielia and its front ends run
against @buildaharness/harness + @buildaharness/runtime by workspace link.
npm install
npm run build:harness && npm run build:runtime # workspace deps the assistant imports
# CLI — the fastest loop. claude-cli backend needs no API key (shells to `claude`).
printf 'what time zone is Tokyo in?\nexit\n' | \
ASSISTANT_LLM_BACKEND=claude-cli node packages/aielia/dist/cli.js
# or, without a build step:
npm run cli --workspace=packages/aielia
# Browser build (chat-ui / the /try page)
npm run dev --workspace=packages/chat-ui # → http://localhost:3010, paste a key in Settings
npm test --workspace=packages/aielia # ~900 tests, must stay green
npm test --workspace=packages/chat-ui
The full 11-layer harness runs client-side every turn — the harness loop itself
makes no LLM calls (it's synchronous state-machine bookkeeping), so the per-turn
cost is one real model call, or zero for a risk-gated one. packages/aielia/README.md
is the design reference; keep it in sync with behaviour changes. Any change to
risk classification, the tool-policy gate, or the approval flow needs a test that
pins the new behaviour.
If you touch a counted number quoted in README.md / README_CN.md / docs/,
run node scripts/gen-stats.mjs (CI runs --check).
Publishing a marketplace component
curl -s -X POST http://localhost:8000/marketplace \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"slug": "my-tool",
"name": "My Tool",
"description": "Does something useful",
"category": "tool",
"icon_emoji": "🔧",
"npm_ref": "@my-scope/my-tool",
"source": "npm",
"node_spec": {"type": "tool_invoke", "tool_id": "my_tool", "data": {"label": "My Tool"}},
"tool_def": {
"tool_ref": "@my-scope/my-tool",
"source": "npm",
"description": "Does something useful",
"input_schema": {"type": "object", "properties": {"input": {"type": "string"}}, "required": ["input"]}
},
"tags": ["category", "keyword"]
}'
Slug rules — kebab-case, [a-z0-9][a-z0-9-]*[a-z0-9], max 80 chars, globally unique.
node_spec.tool_id must equal slug.replace("-", "_").
User-published components start unverified. The @buildaharness verified badge is reserved for the six seed packages.
Adding example flows
Example flows live in flows/. A valid flow must:
- Pass validation against
spec/schema.json - Include
positioncoordinates on all nodes - Have a
descriptionfield - Exercise at least one feature not covered by the existing five flows
- Follow naming:
NN-descriptive-name.json
Adapter test suite
# Main suite (SQLite in-memory, no stack required)
pytest adapter/tests/ -v
pytest adapter/tests/test_maf_adapter.py -v # MAF suite
pytest adapter/tests/test_sso.py -v # SSO/OIDC + SCIM suite
# Harness suite (all infrastructure-free, uses --noconftest)
PYTHONPATH=adapter python3.12 -m pytest adapter/tests/test_harness_p*.py adapter/tests/test_harness_process_concepts.py adapter/tests/test_harness_primitives.py -v --noconftest
# Harness integration + E2E + invariants
PYTHONPATH=adapter python3.12 -m pytest adapter/tests/test_harness_integration_*.py adapter/tests/test_harness_e2e.py adapter/tests/test_harness_invariants.py -v --noconftest
Tests use an in-memory SQLite database via the client fixture in conftest.py. New test files must:
- Use
clientandauth_headersfixtures fromconftest.py - Register a fresh user per test with a unique email address
- Mark all async tests with
@pytest.mark.asyncio
Database migrations
The current migration chain is 0001 → 0011. Add new migrations as 000N_descriptive_name.py in adapter/migrations/versions/:
revision: str = "0012"
down_revision: str | None = "0011"
def upgrade() -> None:
op.create_table("my_table", ...)
def downgrade() -> None:
op.drop_table("my_table")
Use postgresql.UUID and postgresql.JSONB in migrations. In db.py ORM models, use the _UUIDType and _JSONBType wrappers — these fall back to TEXT on SQLite so the CI test suite works without Postgres.
Code of conduct
Be direct. Disagree on specifics, not people. If a decision is documented in an ADR, a closed issue, or the CHANGELOG — open a new issue with new evidence rather than re-litigating the original thread.