Contributing to capy
June 17, 2026 · View on GitHub
Local Setup
Prerequisites
- Go 1.25+ (install)
- C compiler — required for CGO (SQLite FTS5 + sqlite3mc encryption):
- macOS:
xcode-select --install - Debian/Ubuntu:
sudo apt install build-essential - Fedora:
sudo dnf install gcc - No external SQLite library needed — the build uses a
go.modreplace directive to bundle sqlite3mc (SQLite3MultipleCiphers) via the jgiannuzzi/go-sqlite3 fork
- macOS:
- Language runtimes (optional) — for testing the executor against specific languages:
- At minimum:
bash,python3 - Full set:
node/bun,tsx/ts-node,ruby,go,rustc,php,perl,Rscript,elixir
- At minimum:
Clone and Build
git clone https://github.com/serpro69/capy.git
cd capy
make build
This produces a ./capy binary. The -tags fts5 build tag is handled by the Makefile.
Verify
Tests require the CAPY_DB_KEY environment variable (any non-empty string works for tests):
export CAPY_DB_KEY=test-key-for-development
make test # run all tests
make vet # static analysis
make test-race # tests with race detector
Project Structure
cmd/capy/ CLI entry points (serve, hook, setup, doctor, cleanup, checkpoint, encrypt, sweep, which, dbsize)
internal/
adapter/ Platform adapter interface + Claude Code implementation
config/ TOML config loading, project root detection, path resolution
executor/ Polyglot code executor (11 languages, process isolation)
giturl/ Git platform URL detection (shared by hook and server)
hook/ Hook event routing (PreToolUse, PostToolUse, SessionStart, etc.)
platform/ Setup command, doctor diagnostics, routing instructions
sanitize/ Secret stripping (regex-based redaction for indexed content)
security/ Settings parsing, glob matching, command splitting, shell-escape detection
server/ MCP server, 9 tool handlers, stats, lifecycle, snippets, intent search
session/ Claude Code session JSONL parsing, transcript building, chunking, sweep indexing
store/ SQLite FTS5 knowledge base (schema, indexing, chunking, search, cleanup, encryption, migration)
version/ Version variable (set at build via ldflags)
docs/
adr/ Architecture Decision Records (numbered, append-only)
done/ Design docs for completed features (design.md, implementation.md, tasks.md)
wip/ Design docs for in-progress features
architecture.md System architecture overview
For the full architecture, see docs/architecture.md.
Testing
Running Tests
All tests require:
- The
fts5build tag (handled by the Makefile) - The
CAPY_DB_KEYenvironment variable set to any non-empty string
export CAPY_DB_KEY=test-key-for-development # set once per shell session
make test # all tests
make test-race # with race detector
go test -tags fts5 -count=1 ./internal/store/... # single package
go test -tags fts5 -count=1 -run TestSearch ./internal/store/... # single test
go test -tags fts5 -count=1 -v ./internal/hook/... # verbose
Top gotchas:
- Without
-tags fts5you'll get crypticno such module: fts5errors from SQLite. - Without
CAPY_DB_KEYyou'll getCAPY_DB_KEY environment variable is requirederrors from store tests.
Coverage
CGO_ENABLED=1 go test -tags fts5 -coverprofile=cover.out ./...
go tool cover -func=cover.out | tail -1 # total percentage
go tool cover -html=cover.out -o cover.html # visual report
Test Organization
Each package has two levels of tests:
- Unit tests (
*_test.go) — test individual functions in isolation. Use test stubs/mocks. Fast. - Integration tests (
integration_test.go) — test the full pipeline across packages. Use real adapters, real SQLite, real executor. Found ininternal/server/andinternal/hook/.
Key integration test files:
| File | What it covers |
|---|---|
server/integration_test.go | MCP tool handler → executor → store → search round-trips |
hook/integration_test.go | Full hook JSON through Claude Code adapter with real routing |
Benchmarks
Benchmarks are gated behind CAPY_BENCH_RESULTS and never run during go test ./.... They require the same CAPY_DB_KEY as regular tests (see Running Tests above).
Prerequisites:
export CAPY_DB_KEY=test-key-for-development # required — same as tests
go install golang.org/x/perf/cmd/benchstat@latest # optional — only for `make bench-compare`
Running benchmarks after implementing a feature:
# 1. Generate a baseline from main (skip if you already have bench-results/main.*)
git worktree add /tmp/capy-bench-main main
cd /tmp/capy-bench-main && make bench
cp bench-results/main.* /path/to/capy/bench-results/
cd /path/to/capy && git worktree remove /tmp/capy-bench-main
# 2. Run benchmarks on your feature branch
make bench
# 3. Compare
# Branch names with slashes are sanitized: feat/my-thing → feat-my-thing
make bench-compare BASE=main TARGET=feat-my-thing
If you don't have uncommitted changes, a simpler alternative:
git checkout main && make bench && git checkout -
make bench
make bench-compare BASE=main TARGET=feat-my-thing
make bench runs two targets:
| Target | What it measures | Output |
|---|---|---|
make bench-perf | Index throughput, search latency, executor overhead | bench-results/{branch}.txt |
make bench-quality | Retrieval quality (R@K, MRR, NDCG) and context reduction (NIAH) | bench-results/{branch}.json |
Branch names with / are converted to - in filenames (e.g., feat/bench → feat-bench).
Reading results:
# Single report — absolute metrics
go run -tags fts5 ./cmd/qualstat bench-results/main.json
# Two reports — comparison with regression markers
go run -tags fts5 ./cmd/qualstat bench-results/main.json bench-results/feat-my-thing.json
FTS5 warnings during benchmarks: make bench-quality prints many WARN search row iteration failed lines to stderr. These are expected — queries containing FTS5 special characters (quotes, parentheses, OR) fail at the first search layer and fall through to the next. The search fallback chain is working as designed. The benchmark report (failures: 0) is the source of truth, not stderr noise.
Interpreting qualstat output:
~means no change (within epsilon)+0.030means improvement-0.060 !means regression — the!marker appears when a metric drops below its warning threshold- Exit code 0 = clean, exit code 1 = regressions detected
Warning thresholds:
- Zero tolerance: Context Recall, Match-Layer Accuracy — any regression is flagged
- -0.02 tolerance: R@K, MRR, NDCG, Compression — small trade-offs from tuning are acceptable
Adding or modifying fixtures: See benchmark/FIXTURES.md for the full schema, content type reference, match layer guide, needle design rules, and gotchas (negative query authoring, JSON rank ceilings, corpus density). Changing any fixture invalidates existing reports — qualstat refuses to compare across different dataset hashes.
Updating benchmark docs after new results:
# Generate markdown tables from a JSON report
go run -tags fts5 ./cmd/qualstat --markdown bench-results/main.json
Copy the output into benchmark/RESULTS.md, replacing the existing tables. The prose sections (methodology, caveats, known limitations) stay as-is — only the tables change.
Full benchmark methodology and current results: benchmark/RESULTS.md
Local Verification Without Installation
You don't need to install capy globally to test it. Build and run from the repo root.
Testing Hooks Manually
Build the binary and pipe JSON to stdin:
make build
# Test: curl gets intercepted
echo '{"tool_name":"Bash","tool_input":{"command":"curl https://example.com"},"session_id":"test"}' \
| ./capy hook pretooluse | python3 -m json.tool
# Expected: permissionDecision "allow" with updatedInput replacing curl with echo message
# Test: WebFetch gets denied
echo '{"tool_name":"WebFetch","tool_input":{"url":"https://example.com"},"session_id":"test"}' \
| ./capy hook pretooluse | python3 -m json.tool
# Expected: permissionDecision "deny"
# Test: normal bash gets guidance
echo '{"tool_name":"Bash","tool_input":{"command":"ls -la"},"session_id":"test"}' \
| ./capy hook pretooluse | python3 -m json.tool
# Expected: additionalContext with capy guidance (or empty if guidance was already shown this session)
# Test: security deny
echo '{"tool_name":"Bash","tool_input":{"command":"sudo rm -rf /"},"session_id":"test"}' \
| ./capy hook pretooluse --project-dir . | python3 -m json.tool
# Expected: permissionDecision "deny" (if you have a Bash(sudo *) deny rule in .claude/settings.json)
Testing Doctor
./capy doctor
./capy doctor --project-dir /path/to/some/project
Testing With Claude Code (Full E2E)
# Create a scratch project
mkdir /tmp/capy-test && cd /tmp/capy-test
git init
# Setup using your local binary
/path/to/capy/capy setup --binary /path/to/capy/capy
# Verify
/path/to/capy/capy doctor
# Now open this project in Claude Code — hooks and MCP tools should be active
To iterate: rebuild with make build, then restart Claude Code (the MCP server and hooks pick up the new binary automatically since the config points to the absolute path).
Testing the MCP Server Directly
The MCP server uses stdio (JSON-RPC over stdin/stdout). The integration tests cover this path:
go test -tags fts5 -count=1 -v -run TestIntegration ./internal/server/...
Development Workflow
Making Changes
- Find the relevant package in
internal/ - Read existing tests to understand the patterns
- Write or modify code
- Run the package tests:
go test -tags fts5 -count=1 ./internal/<package>/... - Run
make vetfor static analysis - Run
make testfor the full suite - If you changed hook behavior, test manually with piped JSON (see above)
Common Patterns
Tool handlers (server/tool_*.go): Each MCP tool has a handler function that:
- Parses arguments from the request (with input coercion for double-serialized JSON)
- Runs security checks (deny policies, file path checks, shell-escape detection)
- Does the work (execute, index, search, etc.)
- Tracks stats via
s.trackToolResponse() - Returns
*mcp.CallToolResult
Store operations: All database operations use prepared statements cached on the ContentStore struct. New queries need a prepared statement added to store.go:prepareStatements() and closed in store.go:Close().
Write transactions: Use beginImmediate() (no-op DELETE after BEGIN) to acquire SQLite's RESERVED write lock immediately, preventing interleaving between the dedup SELECT and subsequent INSERT/UPDATE.
SQLite WAL checkpoint (see ADR-015, ADR-016, ADR-019): The knowledge DB uses WAL mode with mandatory encryption (sqlite3mc, SQLCipher v4 compat). On Close(), the WAL must be flushed into the main .db file — otherwise git operations corrupt the database. The checkpoint requires exclusive WAL access, which means the database/sql connection pool must be closed before checkpointing. Close() handles this by: (1) closing statements, (2) closing the pool, (3) opening a fresh single connection for PRAGMA wal_checkpoint(TRUNCATE).
WAL/rekey incompatibility (see ADR-020): sqlite3mc does not support PRAGMA rekey in WAL journal mode. capy encrypt's encryptPlain path must switch to DELETE journal mode before rekeying.
Security checks: Bash deny patterns are loaded once at server startup. The matchesAnyBashPattern function uses cached regexes (sync.Map). Shell-escape patterns for non-shell languages are compiled once in init().
Hook routing (hook/pretooluse.go): The main routing function dispatches on canonical tool name. New tool interceptions go here. Guidance uses file-based persistence since hooks run as separate short-lived processes.
Content indexing: All content passes through sanitize.StripSecrets() before hashing and storage. SHA-256 content hashing enables dedup — same content with same label skips re-indexing.
Search pipeline: RRF (Reciprocal Rank Fusion) across Porter + trigram layers, with fuzzy Levenshtein correction on sparse results. Post-processing: per-source diversification, title-match boost, proximity reranking, entity-aware boosting.
Adding a New MCP Tool
- Define the tool schema in
server/tools.go(follow existing patterns for annotations) - Create the handler in a new
server/tool_<name>.gofile - Register it in
server/tools.go:registerTools() - Add the tool name to
platform/routing.go:CapyToolNames - Write unit tests in
server/tool_<name>_test.go - Add an integration test in
server/integration_test.goif the tool interacts with the store
Adding a New Tool Extractor (Session Indexing)
The session parser uses a table-driven ExtractorRegistry (internal/session/tools.go) to decide how each tool's tool_use input appears in indexed transcripts. Three actions:
ActionPromote— tool input becomes part ofAssistantText(survives even on tool-only turns). Used for conversational tools like PAL.ActionEnrich— tool input becomes a metadata line (e.g.,[Read: path/to/file.go]). Only appears on turns that already have text.ActionSkip— tool is omitted from transcript metadata. Default for unregistered tools.
To add a new extractor:
- Write an extract function:
func(input json.RawMessage) string— parse the JSON input, return a human-readable string (empty string = graceful skip) - Register it in
NewDefaultRegistry()intools.gowith the exact tool name and appropriate action - Add a test in
tools_test.gocovering valid input, malformed input, and empty fields
Adding a New Hook Interception
- Add the routing logic to
hook/pretooluse.go:handlePreToolUse() - If it's a new canonical tool, add it to
hook/helpers.go:toolAliasesfor platform variants - Add the tool name to
platform/setup.go:PreToolUseMatcherPatternso the hook is registered - Write tests in
hook/hook_test.go(unit, with test adapter) andhook/integration_test.go(with Claude Code adapter)
Adding Platform Support
- Implement the
adapter.HookAdapterinterface for the new platform - Add tool name aliases to
hook/helpers.go:toolAliases - Add a setup path in
platform/setup.go(e.g.,SetupCodexfor Codex CLI) - Add platform-specific routing instructions if needed
Configuration Reference
capy uses TOML configuration with three-level precedence (lowest to highest):
~/.config/capy/config.toml(global).capy/config.toml(project).capy.toml(project root)
[store]
# path = ".capy/knowledge.db" # optional override; default: ~/.local/share/capy/<project-hash>/knowledge.db
# title_weight = 2.0 # BM25 title column weight
# max_source_bytes = 2097152 # 2 MB hard cap on total content per source
[store.cleanup]
# cold_threshold_days = 30 # durable sources older than this may be evicted
# ephemeral_ttl_hours = 24 # lifetime for ephemeral sources (minimum 1)
# session_ttl_days = 60 # lifetime for session sources (minimum 1)
# auto_prune = false # automatic pruning (not yet implemented)
[store.cache]
# fetch_ttl_hours = 24 # skip re-fetch within this window
[executor]
# timeout = 30 # seconds per execution
# max_output_bytes = 102400 # 100 KB output cap
[server]
# log_level = "info" # "debug", "info", "warn", "error"
Worktree DB resolution (see ADR-026): with a relative (project-scoped)
store.path, a session running in a linked git worktree resolves the DB against
the repository's main worktree, so all worktrees share one committed
.capy/knowledge.db. config.MainWorktreeDir detects this by parsing the
worktree's .git file (no git subprocess; submodules excluded), falling back to
the current directory on any unexpected layout. Absolute store.path and the XDG
default are unaffected. This redirect lives in config.DBProjectDir, which every
store call site routes through, so the CLI commands (capy sweep, cleanup,
checkpoint, dbsize, which) get it too — run from a worktree, they all operate
on the main worktree's DB.
Project-root detection (config.DetectProjectRoot): resolution precedence is
CLAUDE_PROJECT_DIR env → git rev-parse --show-toplevel → walk up for .git/
.capy.toml/.capy → cwd. Claude Code sets CLAUDE_PROJECT_DIR only in the hook
execution environment (the session's launch dir); manual CLI runs don't have it and
fall back to git rev-parse, which returns the current worktree. These usually
agree. They diverge only if the launch dir differs from where the sessions being
swept actually live (e.g. a hook fires with CLAUDE_PROJECT_DIR pointing at the main
checkout while sessions are under the worktree's mangled path) — then capy sweep
looks under the launch dir's session directory, not the worktree's.
Releases
Releases are fully automated. Push a semver tag to trigger the pipeline:
git tag v1.0.0 && git push origin v1.0.0
What happens:
- CI runs vet + tests (same as every push to master)
- Build produces binaries on native runners for 3 platforms (darwin/arm64, linux/amd64, linux/arm64)
- Release creates a GitHub Release with tarballs and SHA256SUMS
- Homebrew updates the
serpro69/homebrew-tapformula (stable releases only)
Pre-release tags (e.g., v1.0.0-rc.1) create pre-release GitHub Releases and skip the Homebrew update.
Code Conventions
- Error handling: Wrap errors with context via
fmt.Errorf("doing X: %w", err). Don't swallow errors silently — at minimum log withslog.Debugorslog.Warn. - Concurrency: Use
sync.Mutexfor mutable state,sync.Oncefor lazy init,sync.Mapfor concurrent caches. No naked goroutines that use shared resources without synchronization. - No
init()side effects: The onlyinit()in the codebase compiles regex patterns (security/shell_escape.go). Keep it that way. - Tests: Use
testify/assertandtestify/require.requirefor preconditions (test stops on failure),assertfor the actual check. - Build tags: Always pass
-tags fts5for anything involving the store or server packages. - Logging: Use
slogexclusively.slog.Infofor operational events,slog.Debugfor internal details,slog.Warnfor recoverable issues.