Awesome Agent Harness

August 30, 2026 · View on GitHub

A curated, implementation-first list of agent harness engineering resources, with GitHub projects as the primary focus.

  • Total entries: 367
  • GitHub entries: 333 (90.7%)
  • GitHub in project categories (excluding readings): 328/328 (100.0%)
  • Categories: 9
  • Last verified: 2026-08-31
  • Language: English | 中文

Contents

Category Overview

CategoryEntries
Harness Architecture & Orchestration64
Context & Working-State Engineering29
Execution Substrates & Sandboxing27
Protocols, Tool Interfaces & Agent Contracts42
Evaluation Harnesses & Benchmarks29
Observability & Reliability Operations21
Guardrails, Security & Governance27
Reference Harness Implementations89
Essential Readings & Ecosystem Maps39

Catalog

Notes:

  • Stars are rendered as badges from snapshot values.
  • Repository update dates are tracked in data/projects.yaml and validation reports.
  • Entries are sorted by stars (descending) within each category.

Harness Architecture & Orchestration

ProjectLinkStarsTagsSummary
SuperpowersGitHubstarskills, workflow, cross-agentCross-agent software development methodology built from composable skills, mandatory workflows, worktrees, planning, TDD, review, and subagent execution.
ECCGitHubstarcross-harness, hooks, skillsCross-harness operator system combining skills, hooks, memory optimization, security scanning, and validation workflows for agentic work.
Matt Pocock SkillsGitHubstarskills, coding-agents, workflowsComposable engineering skills for Claude Code, Codex, and other coding agents, covering alignment, shared language, ADRs, specs, tests, and code review loops.
gstackGitHubstarskills, qa, releaseClaude Code and cross-agent skill stack that turns product planning, architecture review, QA, security, release, and retrospectives into repeatable agent workflows.
Addy's Agent SkillsGitHubstarskills, quality-gates, coding-agentsProduction-grade engineering skills for coding agents that package lifecycle workflows, quality gates, reviews, testing, debugging, security, and release practices.
DeerFlowGitHubstarlong-horizon, memory, subagentsLong-horizon super-agent harness integrating memory, tools, subagents, and sandboxes.
RufloGitHubstarmulti-agent, swarm, mcpMulti-agent orchestration platform for Claude Code with swarms, persistent memory, federation, plugins, and MCP hooks.
oh-my-openagentGitHubstarmulti-harness, team-mode, skillsMulti-harness agent OS for OpenCode, Codex, Claude Code, and other coding agents with team-mode orchestration, background agents, MCPs, and skills.
AutoGenGitHubstarmulti-agent, orchestration, frameworkProgramming framework for agentic AI with multi-agent interaction and orchestration.
CrewAIGitHubstarmulti-agent, workflows, control-planeMulti-agent automation framework with production Flows, autonomous Crews, event-driven control, tracing, guardrails, memory, and human review hooks.
OrcaGitHubstarcontrol-plane, worktrees, cross-agentCross-agent desktop and mobile control plane for running CLI coding agents in isolated worktrees, monitoring their state, reviewing changes, and steering remote or local sessions.
AgnoGitHubstarscale, runtime, managementAgent software runtime focused on running and managing agentic systems at scale.
LangGraphGitHubstargraph, workflow, runtimeGraph-based runtime for resilient stateful agents and deterministic workflow control.
Claude Code RouterGitHubstarcontrol-plane, routing, observabilityLocal model gateway and control plane for coding agents with provider profiles, conditional routing, credential pools, retries, fallbacks, tool extensions, quotas, and request traces.
ChatDevGitHubstarmulti-agent, orchestration, workflowsZero-code multi-agent orchestration platform for defining agents, workflows, and tasks, evolved from a virtual software-company harness.
HerdrGitHubstarterminal-runtime, persistent-sessions, cross-agentPersistent terminal runtime for coding agents with reconnectable sessions, blocked and idle detection, agent-driven orchestration APIs, and plugin-extensible workflows.
AionUiGitHubstarmulti-agent, control-plane, schedulingCross-platform cowork control plane with a built-in agent engine, external CLI-agent orchestration, team mode, per-agent approvals, remote channels, MCP management, and scheduled runs.
BuzzGitHubstarcollaboration, event-log, auditSelf-hostable human-agent workspace built on a signed event log, with scoped identities, audit trails, agent-first CLI and ACP bridges, workflows, Git events, search, and approvals.
OpenAI Agents SDK (Python)GitHubstarsdk, handoff, workflowsLightweight framework for multi-agent workflows, handoffs, and production patterns.
deepagentsGitHubstarruntime, orchestration, long-runningOpen-source harness for long-running, tool-using agents with planning and subagent patterns.
Semantic KernelGitHubstarenterprise, orchestration, pluginsEnterprise-grade agentic application framework with orchestration and plugin patterns.
SymphonyGitHubstarorchestration, control-plane, workflowsTicket-driven orchestration layer that turns project work into isolated autonomous implementation runs.
ArchonGitHubstarworkflow-engine, worktrees, validationWorkflow engine for AI coding agents with YAML-defined phases, isolated worktrees, and validation gates.
Go MicroGitHubstargo, service-runtime, durable-flowsGo agent harness and service framework combining model loops, durable memory, tools, planning, delegation, guardrails, service discovery, MCP, A2A, and durable flows.
Google ADK (Python)GitHubstartoolkit, deployment, evaluationCode-first toolkit to build, evaluate, and deploy advanced AI agents.
PydanticAIGitHubstarpython, typing, schemaType-safe Python framework for agents with strong schema contracts and tooling.
elizaOSGitHubstaragent-os, plugins, benchmarksExtensible agent runtime and operating system with CLI scaffolding, agent loop, plugins, memory/state primitives, dashboards, connectors, and benchmark suites.
Gas TownGitHubstarmulti-agent, workspaces, coordinationMulti-agent workspace manager for coordinating coding agents with persistent work tracking, git-backed hooks, handoffs, supervision, and merge queues.
CAMELGitHubstarmulti-agent, scalability, benchmarksMulti-agent framework for large-scale agent systems with communication, stateful memory, environments, benchmarks, and coordination primitives.
QMGitHubstarmulti-user, durable-sandbox, policyMultiplayer organizational agent harness with per-person and per-room scopes, durable sandboxes, memory, permissions, schedules, policy controls, and swappable coding-agent runtimes.
CloudCLIGitHubstarcontrol-plane, sessions, coding-agentsWeb and mobile control surface for Claude Code, OpenCode, Cursor CLI, Codex, and Gemini CLI with project, terminal, file, Git, session, and plugin management.
Microsoft Agent FrameworkGitHubstarmulti-agent, workflows, observabilityMulti-language framework for building, orchestrating, and deploying AI agents with graph workflows and observability.
HiveGitHubstarharness, orchestration, runtimeOutcome-driven agent runtime harness with explicit control loops and orchestration blocks.
VoltAgentGitHubstartypescript, platform, runtimeTypeScript agent engineering platform built around open runtime abstractions.
OmnigentGitHubstarmeta-harness, policies, sandboxesOpen-source meta-harness that orchestrates Claude Code, Codex, Cursor, Pi, and custom agents with shared sessions, cloud sandboxes, policies, and collaboration.
PraisonAIGitHubstarmulti-agent, workflow, memoryMulti-agent workforce framework with autonomous planning, execution, memory, RAG, dashboards, and multi-provider model support.
mcp-agentGitHubstarmcp, runtime, workflowPractical agent framework centered on MCP tool ecosystems and workflow composition.
Claude SquadGitHubstarterminal, sessions, worktreesTerminal control plane for managing multiple Claude Code, Codex, Gemini, OpenCode, Amp, and other agent sessions with isolated git workspaces and review flows.
FlueGitHubstartypescript, headless, sandboxTypeScript harness framework for building headless agents with sessions, tools, skills, and pluggable sandboxes.
YaoGitHubstarsingle-binary, runtime, autonomousSingle-binary runtime for defining and running autonomous agents.
Agent SquadGitHubstarrouting, multi-agent, contextMulti-agent orchestration framework that routes requests, preserves conversation context, supports Python/TypeScript, and coordinates specialist agents.
Strands AgentsGitHubstarsdk, mcp, toolsModel-driven agent SDK and monorepo with Python/TypeScript agent loops, provider adapters, tools, MCP integration, multi-agent systems, and streaming.
Open Multi-AgentGitHubstarmulti-agent, dag, tracingTypeScript-native multi-agent orchestrator that turns goals into task DAGs with parallel execution, MCP integration, and live tracing.
AIOSGitHubstaragent-os, kernel, runtimeAI Agent Operating System with a kernel and SDK for scheduling, context management, memory, storage, tools, deployment modes, and computer-use sandboxing.
NexentGitHubstarzero-code, control-plane, multi-agentZero-code agent platform built around harness engineering principles, unifying tools, skills, memory, orchestration, constraints, feedback loops, and control planes.
Cloudflare AgentsGitHubstarplatform, deployment, runtimePlatform runtime for building and deploying agents with production infrastructure primitives.
Embabel Agent FrameworkGitHubstarjvm, planning, typed-flowsJVM agent framework for typed agentic flows with goals, actions, conditions, dynamic planning, platform modes, and testability.
cascadeflowGitHubstarruntime-policy, model-routing, budget-gatesIn-process agent harness for per-step model routing, tool budget gates, runtime policy actions, and auditable control inside agent loops.
OpenAI Agents SDK (JS/TS)GitHubstartypescript, workflows, sandbox-agentsJavaScript/TypeScript framework for multi-agent workflows with handoffs, tools, guardrails, sessions, tracing, and sandbox agents.
Docker AgentGitHubstardocker, runtime, containerAgent builder and runtime stack emphasizing container-native execution.
NeMo Agent ToolkitGitHubstarmulti-agent, optimization, toolkitOpen toolkit for connecting and optimizing teams of AI agents.
Apache BurrGitHubstarstate-machine, persistence, tracingState-machine framework for decision-making agents and LLM apps with persistence, telemetry UI, tracing, and framework-agnostic execution.
ShannonGitHubstarmulti-agent, temporal, sandboxProduction-oriented multi-agent orchestration framework with Temporal workflows, WASI sandboxing, OPA approvals, token budgets, and OpenTelemetry tracing.
AXGitHubstardistributed-runtime, isolated-actors, policyGoogle's distributed agent runtime for coordinating agent loops, isolated actors, event logging, auditing, and policy.
tRPC-Agent-GoGitHubstargo, graph-workflows, observabilityGo framework for production agent systems with graph workflows, tools, memory, evaluation, protocols, and OpenTelemetry observability.
ScionGitHubstarmulti-agent, containers, orchestrationExperimental multi-agent orchestration testbed that runs isolated agent harnesses in containers, worktrees, and remote runtimes.
deepagentsjsGitHubstartypescript, langgraph, subagentsTypeScript agent harness with built-in planning, filesystem tools, subagents, and LangGraph-native runtime hooks.
oh-my-agentGitHubstarmulti-agent, skills, cross-runtimePortable multi-agent harness that projects shared agents, skills, workflows, and rules into multiple coding-agent runtimes.
LiteLLM Agent Control PlaneGitHubstarcontrol-plane, sessions, runtimeAgent control plane that provides one API and UI across OpenCode, Hermes, Claude Managed Agents, Cursor Agents API, DeepAgents, and OpenClaw runtimes.
ChorusGitHubstarai-dlc, permissions, task-stateAI-human collaboration harness for session lifecycle, task state, sub-agent orchestration, observability, and recovery.
Pydantic AI HarnessGitHubstarcapabilities, hooks, pydanticOfficial Pydantic AI capability library for composing tools, lifecycle hooks, instructions, and model settings into reusable agent harnesses.
WaterGitHubstarpython, framework, approval-gatesPython agent harness framework for orchestration, resilience, observability, guardrails, approval gates, sandboxing, and deployment.
OmniCoreAgentGitHubstarpython, mcp, servingPython production harness with model loop, tools, MCP, memory, workspace files, guardrails, events, subagents, background tasks, and REST/SSE serving.
hankweaveGitHubstarlong-horizon, runtime, checkpointsHeadless-first long-horizon runtime that orchestrates existing agent harnesses with sentinels, loops, checkpoints, and event journals.

Context & Working-State Engineering

ProjectLinkStarsTagsSummary
GraphifyGitHubstarknowledge-graph, skills, codexCross-agent skill and CLI that turns code, docs, schemas, and media into a queryable knowledge graph with Codex and Claude Code hooks.
claude-memGitHubstarmemory, context, sessionPlugin-style memory layer that captures session history and reinjects relevant context into future coding runs.
CodeGraphGitHubstarknowledge-graph, mcp, code-contextLocal code knowledge-graph MCP server and installer that wires semantic code intelligence into Claude Code, Codex, Gemini, Cursor, OpenCode, and other agents.
HeadroomGitHubstarcontext-compression, tool-output, mcpContext-compression layer for agent tool outputs, logs, files, RAG chunks, and conversations, available as a library, proxy, agent wrapper, and MCP server.
Mem0GitHubstarmemory, state-management, retrievalUniversal memory layer for AI agents with user, session, and agent state, hybrid retrieval, SDKs, and self-hosted deployment.
MemPalaceGitHubstarmemory, local-first, mcpLocal-first verbatim memory layer with scoped semantic retrieval, pluggable storage backends, conversation mining, MCP integration, and reproducible LongMemEval results.
OpenVikingGitHubstarcontext-database, memory, retrievalContext database for agents that unifies memory, resources, and skills in a virtual filesystem with tiered loading, recursive retrieval, observable trajectories, and session-to-memory extraction.
code-review-graphGitHubstarcode-context, mcp, incremental-indexingLocal-first code intelligence graph that incrementally maps dependencies and exposes targeted review context, blast-radius analysis, and risk-scored CI gates through MCP and CLI.
agentmemoryGitHubstarmemory, mcp, hooksPersistent memory server for coding agents using hooks, MCP/REST integration, hybrid search, and shared session recall.
BeadsGitHubstarmemory, issue-tracking, work-stateAgent-optimized distributed issue tracker that stores long-horizon coding work as dependency-aware graph state with memory recall and multi-branch sync.
planning-with-filesGitHubstarplanning, skills, persistenceSkill package for persistent file-based planning in coding-agent workflows.
TencentDB Agent MemoryGitHubstarmemory, context-offloading, openclawLocal agent memory plugin combining symbolic short-term state, layered long-term memory, traceability, and OpenClaw/Hermes integrations.
HindsightGitHubstarmemory, learning, mcpProduction agent memory system with retain, recall, and reflect operations, isolated memory banks, evidence-backed observations, coding-agent integrations, MCP, and memory-defense policies.
Context ModeGitHubstarcontext, mcp, sessionMCP context optimization server that sandboxes tool output, indexes session events, and restores continuity across agent compactions.
Agent Skills for Context EngineeringGitHubstarskills, context, productionLarge skill library oriented around context engineering and production agents.
SkillOptGitHubstarskills, optimization, validation-gatesMicrosoft optimizer that trains reusable natural-language agent skills through trajectory edits, validation gates, and deployable skill artifacts.
memUGitHubstarmemory, context, retrievalMemory harness for AI agents that turns raw workspace data into structured, queryable context for agent retrieval.
TrellisGitHubstarspecs, memory, workflowMulti-platform coding-agent workflow framework with task context, project memory, and spec injection.
EverOSGitHubstarmemory, local-first, skillsPortable self-evolving memory runtime for agents, persisting readable Markdown with local SQLite/LanceDB indexes and reusable cases and skills.
Context-Engineering HandbookGitHubstarcontext-engineering, handbook, practicesFirst-principles handbook focused on practical context engineering for agent systems.
CCPMGitHubstarplanning, github-issues, parallel-executionSpec-driven project-manager skill that turns PRDs and GitHub issues into persistent context and parallel agent execution.
EngramGitHubstarmemory, mcp, single-binaryAgent-agnostic persistent memory system packaged as a Go binary with SQLite/FTS5, MCP server, HTTP API, CLI, TUI, and sync workflows.
ByteRover CLIGitHubstarcontext-memory, cli, mcpPortable context-memory CLI for coding agents with a context tree, review workflow, versioning, built-in tools, cloud sync, and MCP integration.
AcontextGitHubstarskills, memory, progressive-disclosureSkill-memory layer that distills agent runs into inspectable skill files and recalls them through agent-controlled tools.
Awesome Context EngineeringGitHubstarawesome-list, context, surveySurvey-style list for context engineering resources and frameworks.
agentic-stackGitHubstarcross-harness, memory, skillsPortable memory, skills, protocols, and dashboard layer that keeps state across multiple coding-agent harnesses.
context-spaceGitHubstarcontext, infrastructure, mcpInfrastructure project focused on context engineering building blocks and MCP-centric integrations.
MemorixGitHubstarmemory, mcp, cross-agentLocal-first cross-agent memory control plane with MCP support, workspace sync, sessions, and orchestration state.
sd0x-dev-flowGitHubstarhooks, state-machine, claude-codeClaude Code harness layer with hook-enforced dual review, durable state-machine gates, context-compaction recovery, and fail-closed safety.

Execution Substrates & Sandboxing

ProjectLinkStarsTagsSummary
DaytonaGitHubstarsandbox, execution, infraSecure and elastic sandbox infrastructure for running AI-generated code with file, Git, LSP, and execution APIs.
agent-browserGitHubstarbrowser, automation, cliNative browser automation CLI for AI agents with accessibility snapshots, semantic locators, screenshots, streaming, and batch execution.
CUAGitHubstarcomputer-use, sandbox, infraInfrastructure stack for computer-use agents with sandbox, SDK, and benchmark support.
Browser HarnessGitHubstarbrowser, cdp, self-healingThin editable CDP harness that connects LLMs directly to real browsers and lets agents extend helpers in flight.
OpenSandboxGitHubstarsandbox, security, runtimeSecure and extensible sandbox runtime built for agent workloads.
E2BGitHubstarcloud-sandbox, execution, enterpriseSecure cloud environments with real tools for production-grade agent execution.
InsForgeGitHubstarbackend, mcp, coding-agentsBackend substrate for agentic coding that exposes database, auth, storage, compute, hosting, logs, and model-gateway operations through MCP and CLI/Skills.
CubeSandboxGitHubstarmicrovm, sandbox, e2b-compatibleMicroVM-based sandbox service for AI agents with sub-60ms startup, E2B-compatible APIs, and hardware-level isolation.
OpenShellGitHubstarsandbox, policy, runtimeSafe private runtime for autonomous agents with sandbox lifecycle control and declarative filesystem, network, process, and inference policies.
MicrosandboxGitHubstarsandbox, vm, mcpRootless local VM sandbox runtime with SDKs, detached long-running sessions, agent skills, and MCP server integration.
SandcastleGitHubstarsandbox, typescript, branch-strategyTypeScript library for orchestrating coding agents inside isolated sandboxes with configurable branch strategies.
agent-infra sandboxGitHubstarall-in-one, browser, shellAll-in-one sandbox combining browser, shell, files, MCP, and IDE server.
Judge0GitHubstarcode-execution, sandbox, backendScalable sandboxed code execution system usable as an agent execution backend.
Agent SandboxGitHubstarkubernetes, sandbox, statefulKubernetes-native sandbox control plane for isolated, stateful agent runtimes with stable identity, persistence, and warm-pool support.
stakpak/agentGitHubstaralways-on, autonomous, opsAlways-on open agent that runs on your machines with autonomous operational loops.
Sandbox AgentGitHubstarsandbox, coding-agents, session-schemaHTTP/SSE control server for running coding agents inside sandboxes with normalized sessions, permissions, event streaming, and replay.
E2B Desktop SandboxGitHubstardesktop, sandbox, computer-useSecure virtual desktop sandbox for computer-use agents with SDK control and screen streaming.
OSS-Fuzz GenGitHubstarfuzzing, security, executionLLM-powered fuzzing workflows integrated with controlled execution contexts.
AgentBay SDKGitHubstarcloud-sandbox, computer-use, sdkCloud sandbox SDK for agents spanning browser, desktop, mobile, and code execution environments.
TensorlakeGitHubstarmicrovm, sandbox, orchestrationServerless runtime for agent sandboxes with MicroVM isolation, snapshots, suspend-resume, and background orchestration.
AgentScope RuntimeGitHubstarruntime, sandbox, deploymentProduction runtime for agent apps with secure tool sandboxes, deployment APIs, observability, and state services.
SWE-ReXGitHubstarsandbox, execution, coding-agentSandboxed execution infrastructure for AI coding agents at local and cloud scale.
sandboxed.shGitHubstarself-hosted, isolation, orchestratorSelf-hosted orchestrator running coding agents inside isolated Linux workspaces.
CapsuleGitHubstarwasm, sandbox, task-runtimeDurable runtime that coordinates agent tasks inside isolated WebAssembly sandboxes with retries and lifecycle tracking.
agentboxGitHubstarsandbox, coding-agents, network-policyLocked-down local sandbox for AI coding agents with scoped filesystem access, egress policy, secret injection, firewalling, and persistent agent state.
HexAgentGitHubstarcomputer-layer, sandbox, runtimeAgent harness that separates the runtime from the computer it operates on through local, VM, and cloud sandbox backends.
terminal-bench-envGitHubstarterminal, benchmark-env, sandboxEnvironment layer for terminal-agent benchmark execution.

Protocols, Tool Interfaces & Agent Contracts

ProjectLinkStarsTagsSummary
Anthropic Agent SkillsGitHubstarskills, spec, claudeOfficial Agent Skills repository containing the skills specification, templates, and reference skill implementations for Claude.
GitHub Spec KitGitHubstarspec-driven, workflows, toolingToolkit for spec-driven development to guide deterministic agent execution.
MCP ServersGitHubstarmcp, servers, implementationsOfficial collection of MCP server implementations across tools and domains.
Chrome DevTools MCPGitHubstarmcp, browser, devtoolsOfficial MCP server that gives coding agents Chrome DevTools access for reliable browser automation, debugging, and performance analysis.
AAS CoreGitHubstarskills, control-plane, validationLocal agent-first control plane for discovering, selecting, validating, and planning reproducible skill stacks through a read-only MCP interface, manifests, plugins, and host-specific installers.
Awesome CopilotGitHubstargithub-copilot, skills, hooksGitHub-managed collection of Copilot agents, instructions, skills, hooks, workflows, plugins, tools, and machine-readable listings.
Playwright MCPGitHubstarmcp, browser, playwrightOfficial Playwright MCP server giving agents structured accessibility snapshots and deterministic browser automation tools.
Claude Code Plugins DirectoryGitHubstarplugins, claude-code, marketplaceAnthropic-managed Claude Code plugin marketplace defining plugin manifests, MCP configuration, commands, agents, skills, and submission quality gates.
GitHub MCP ServerGitHubstarmcp, github, officialGitHub's official MCP server connecting agents to repositories, issues, pull requests, CI workflows, releases, and code analysis tools.
OpenAI Codex Plugin for Claude CodeGitHubstarplugin, codex, cross-agentOfficial Claude Code plugin that exposes Codex review, adversarial review, delegation, status, result, and cancel commands inside another coding-agent harness.
Vercel Agent SkillsGitHubstarskills, vercel, officialVercel's official Agent Skills collection packaging reusable instructions and scripts in the open Agent Skills format for coding-agent workflows.
skillsGitHubstarskills, cli, cross-agentCLI for installing, using, finding, updating, and initializing Agent Skills across OpenCode, Claude Code, Codex, Cursor, and other coding agents.
SerenaGitHubstarmcp, coding-agents, semantic-toolsMCP toolkit that gives coding agents IDE-like semantic retrieval, editing, refactoring, debugging, and memory tools.
DESIGN.mdGitHubstarspec, design-contract, cliFormat specification and CLI for giving coding agents persistent, machine-readable design tokens and human-readable design rationale.
FastMCPGitHubstarmcp, python, frameworkPython framework for building MCP servers and clients with generated schemas, validation, documentation, production deployment patterns, and governance hooks.
Agent Skills SpecificationGitHubstarskills, spec, progressive-disclosureOpen specification and documentation for packaging reusable agent capabilities, workflows, scripts, references, and assets behind progressive disclosure.
MCP Python SDKGitHubstarmcp, python, sdkOfficial Python implementation of MCP for building clients and servers that expose tools, resources, prompts, protocol lifecycle events, and standard transports.
AGENTS.mdGitHubstarspec, agent-file, instructionsOpen format for repository-local instructions that coding agents can follow.
Google Agent SkillsGitHubstarskills, google-cloud, officialOfficial Google Agent Skills repository for Google products and technologies, including Agent Platform and Google Cloud workflows.
MCP Toolbox for DatabasesGitHubstarmcp, databases, googleGoogle open-source MCP server and custom tools framework for connecting agents, IDEs, and applications to enterprise databases.
MCP TypeScript SDKGitHubstarmcp, typescript, sdkOfficial TypeScript MCP SDK with server and client packages, transports, auth helpers, middleware adapters, and runnable examples.
Hugging Face SkillsGitHubstarskills, hugging-face, officialOfficial Hugging Face Agent Skills collection for packaging Hub, dataset, training, evaluation, Spaces, and tooling workflows across major coding agents.
MCP InspectorGitHubstarmcp, debugging, testingOfficial developer tool for testing and debugging MCP servers through an interactive client UI and protocol bridge.
mcp-useGitHubstarmcp, framework, appsFull-stack MCP framework for building MCP apps for ChatGPT and Claude and MCP servers for AI agents.
AWS MCP ServersGitHubstarmcp, aws, serversAWS Labs suite of MCP servers exposing AWS documentation, cloud workflows, infrastructure operations, and service tooling to MCP clients.
Model Context ProtocolGitHubstarmcp, protocol, interoperabilityCore specification and docs for MCP-based tool and context interoperability.
MCP RegistryGitHubstarmcp, registry, discoveryOfficial MCP registry service and publisher tooling that lets clients discover MCP servers through a shared server catalog.
.NET Agent SkillsGitHubstarskills, dotnet, officialMicrosoft .NET team's curated plugins, skills, custom agents, and evaluation dashboard for grounding coding agents on .NET and C# workflows.
SkillHubGitHubstarskills, registry, governanceSelf-hosted enterprise agent skill registry with package publishing, versioning, discovery, namespaces, RBAC, reviews, and audit logs.
Builder.io Agent SkillsGitHubstarskills, review, planningComposable coding-agent skills for visual planning, visual recaps, watchdog reviews, plan arbitration, and autonomous execution discipline.
Agent Client ProtocolGitHubstaracp, protocol, coding-agentsOpen protocol that standardizes communication between code editors and coding agents.
directories (rules and MCP indexes)GitHubstardirectories, mcp, rulesCurated directories of agent rules and MCP servers for tool discovery.
AtmosphereGitHubstarjvm, multi-protocol, governanceJVM runtime for streaming governable AI agents across MCP, A2A, AG-UI, and browser-facing transports.
LangChain MCP AdaptersGitHubstarmcp, adapters, integrationAdapters connecting LangChain components with MCP servers.
Microsoft MCP ServersGitHubstarmcp, enterprise, serversMicrosoft's official MCP server catalog for enterprise data and tools.
ACPXGitHubstaracp, client, sessionsHeadless CLI client for stateful Agent Client Protocol sessions.
Microsoft Agent SkillsGitHubstarskills, mcp, officialMicrosoft-maintained skills, custom agents, AGENTS.md templates, and MCP configurations for Azure SDK and Microsoft AI Foundry coding-agent workflows.
GitAgentProtocolGitHubstarstandard, git-native, workflowsGit-native, framework-agnostic standard for defining agents, skills, workflows, tools, and runtime memory in repositories.
OpenGAPGitHubstargit-native, cli, agent-contractReference CLI for the Git Agent Protocol, turning repository files into portable agent manifests, rules, skills, workflows, tools, memory, hooks, and compliance contracts.
Microsoft Learn MCPGitHubstarmcp, docs, groundingMCP server and CLI for grounding agents with Microsoft documentation sources.
IBM MCPGitHubstarmcp, clients, toolingIBM collection of MCP servers, clients, and developer tooling.
AGENT.mdGitHubstarstandard, agent-file, interoperabilityStandardized machine-readable file format for agentic coding tools.

Evaluation Harnesses & Benchmarks

ProjectLinkStarsTagsSummary
PromptfooGitHubstareval, red-team, ciConfig-driven prompt/agent/RAG testing, comparison, and red-team evaluation tool.
DeepEvalGitHubstarevaluation, framework, testingLLM evaluation framework supporting agent and workflow quality testing.
RAGASGitHubstarrag, metrics, evaluationOpen evaluation toolkit for LLM and RAG quality metrics.
lm-evaluation-harnessGitHubstarbenchmark, harness, llmPopular benchmark harness for consistent LLM evaluation across tasks.
Agent Reinforcement TrainerGitHubstarrl, training, evaluationAgent reinforcement-training harness for multi-step agents with GRPO, trajectory capture, reward assignment, training loops, and LangGraph/MCP examples.
SWE-benchGitHubstarbenchmark, swe, evaluationStandard benchmark for evaluating issue-fixing software engineering agents.
HarborGitHubstarevaluation, harness, rl-envFramework for running agent evaluations and constructing RL-style environments.
verifiersGitHubstarverifier, rl, evaluationLibrary for RL environments and verifier-based evaluation loops.
AgentBenchGitHubstarbenchmark, cross-domain, agentCross-environment benchmark for evaluating LLM agents as tool-using systems.
LangWatchGitHubstarsimulation, evaluation, testingEnd-to-end platform for agent simulations, evaluation loops, and production testing.
EvalScopeGitHubstarbenchmark, framework, llmCustomizable framework for large-model benchmarking and performance evaluation.
TEAQL Agent KitGitHubstarcoding-agent, evaluation, auditabilityControlled evaluation environment for coding agents on auditable business software tasks with API adherence, framework discipline, self-repair, and human-gate metrics.
tau2-benchGitHubstartool-use, interaction, benchmarkTool-agent-user interaction benchmark emphasizing multi-step execution quality.
WebArenaGitHubstarweb-agent, benchmark, environmentSelf-hostable web environment and evaluation harness for autonomous web agents with reproducible end-to-end tasks.
Meta-HarnessGitHubstarharness-search, optimization, terminal-benchFramework for automated search over task-specific model harnesses, with reference experiments for memory systems and terminal-agent scaffolds.
NeMo GymGitHubstarrl-env, training, evaluationToolkit for building RL environments suitable for LLM/agent training and eval.
TheAgentCompanyGitHubstarbenchmark, workplace, multi-stepAgent benchmark with simulated software-company tasks for evaluating multi-step workplace autonomy.
Claw-EvalGitHubstarbenchmark, trajectory, safetyEvaluation harness and benchmark for autonomous agents with human-verified tasks, trajectory auditing, and completion, safety, and robustness rubrics.
Inspect EvalsGitHubstarinspect, eval-suite, reproducibilityEvaluation suite collection for Inspect AI workflows.
ClawBenchGitHubstarbrowser-agent, benchmark, recordingBrowser-agent benchmark with live-site tasks, isolated containers, five-layer recording, and agentic scoring.
Terminal-BenchGitHubstarterminal, benchmark, long-horizonTerminal-native benchmark suite for long-horizon, verification-heavy agent tasks.
auto-harnessGitHubstaroptimization, regression, evalsBenchmark-gated optimization loop that mines failures, edits agent code, and guards against regressions overnight.
WildClawBenchGitHubstarbenchmark, harness-comparison, multimodalIn-the-wild benchmark that compares multiple agent harnesses on end-to-end multimodal, coding, safety, and productivity tasks inside a live OpenClaw environment.
SWE-Bench ProGitHubstarswe, benchmark, long-horizonLong-horizon software-engineering benchmark with reproducible Docker-based evaluation for issue-driven coding agents.
Agent EvaluationGitHubstarevaluation, testing, ciAWS framework for testing virtual agents with evaluator-driven multi-turn conversations, hooks, and CI-friendly workflows.
WorkArenaGitHubstarbrowser, benchmark, enterpriseBrowser benchmark for practical enterprise-like knowledge work tasks.
OpenHands BenchmarksGitHubstaropenhands, eval, harnessEvaluation harness and benchmark definitions for OpenHands systems.
WebArena-VerifiedGitHubstarweb-agent, benchmark, deterministicVerified web-agent benchmark with deterministic evaluators.
HarnessBenchGitHubstarharness-comparison, browser-agent, benchmarkBenchmark for comparing agent harnesses on the same everyday web tasks with fixed models and per-harness containers.

Observability & Reliability Operations

ProjectLinkStarsTagsSummary
LangfuseGitHubstarllmops, tracing, metricsOpen-source LLM engineering platform for traces, metrics, prompts, and evals.
MLflowGitHubstarplatform, monitoring, evaluationBroad AI engineering platform with monitoring and evaluation support for agents.
OpikGitHubstarmonitoring, eval, tracingEnd-to-end debug/eval/monitoring stack for LLM apps and agent workflows.
RagaAI CatalystGitHubstaragentops, analytics, monitoringAgent observability and monitoring framework with timeline and graph analytics.
TensorZeroGitHubstarllmops, gateway, optimizationOpen LLMOps stack unifying gateway, observability, evaluation, and optimization.
Arize PhoenixGitHubstarobservability, tracing, evaluationOpen platform for AI observability, tracing, and evaluation analytics.
OpenLLMetryGitHubstaropentelemetry, instrumentation, tracingOpenTelemetry-based instrumentation for GenAI and LLM applications.
HeliconeGitHubstarmonitoring, traffic, productionLightweight platform for monitoring and evaluating LLM traffic in production.
AgentOps SDKGitHubstaragentops, monitoring, costMonitoring and benchmarking SDK for agent workflows with cost and trace tracking.
Coze LoopGitHubstarevaluation, tracing, monitoringFull-lifecycle agent optimization platform combining prompt development, automated evaluation, experiment management, trace collection, execution inspection, and production monitoring.
AgentaGitHubstarllmops, evaluation, observabilityOpen-source LLMOps platform combining prompt management, evaluation, human and automated feedback, and tracing for production LLM applications.
LatitudeGitHubstarplatform, eval, observabilityOpen-source agent engineering platform with eval and observability capabilities.
TruLensGitHubstarevaluation, tracing, agentopsEvaluation and tracking platform for LLM apps and AI agents with traces, tool-call capture, and agentic evaluators.
LaminarGitHubstarobservability, tracing, evalsAgent-focused observability stack with tracing, evaluation runs, monitoring, and dashboards.
DesloppifyGitHubstarquality-gates, codebase-health, ciAgent-facing codebase quality harness with scans, scoring, LLM review, prioritized fix loops, persistent state, and CI gates.
OpenLITGitHubstaropentelemetry, observability, evaluationsOpenTelemetry-native AI engineering platform for LLM and coding-agent observability, evaluations, guardrails, prompt management, secrets, and dashboards.
claude-code-reverseGitHubstartrace, visualization, debuggingTooling to visualize and inspect Claude Code LLM interaction traces.
Future AGIGitHubstarobservability, evaluation, guardrailsSelf-hostable platform that closes the loop across agent tracing, evaluation, simulation, guardrails, and gateway operations.
OpenInferenceGitHubstarspec, instrumentation, observabilityOpen instrumentation specification and tooling for AI observability.
JudgevalGitHubstartracing, agent-judges, monitoringAgent improvement SDK with OpenTelemetry tracing, agent judges, online monitoring, replay evaluations, CLI workflows, and MCP access to traces and behaviors.
Agentic Harness EngineeringGitHubstarharness-optimization, trace-analysis, terminal-benchObservability-driven system for evolving coding-agent harness components through evaluate, analyze, and improve loops.

Guardrails, Security & Governance

ProjectLinkStarsTagsSummary
OmniRouteGitHubstargateway, routing, guardrailsLocal-first AI gateway for coding agents with quota-aware routing, provider fallback, credential pools, circuit breakers, usage controls, compression, MCP/A2A interfaces, and guardrails.
LiteLLMGitHubstargateway, proxy, guardrailsUnified LLM gateway/proxy with cost tracking, load balancing, and guardrails.
KongGitHubstargateway, policy, infraAPI and AI gateway infrastructure useful for policy enforcement in agent systems.
ParlantGitHubstarinteraction-control, guardrails, customer-agentsInteraction-control harness for customer-facing agents focused on consistent, predictable, and governed LLM behavior.
SkillSpectorGitHubstarskill-security, static-analysis, risk-scoringSecurity scanner for AI agent skills with static and optional LLM analysis, vulnerability patterns, risk scoring, and CI-friendly report formats.
Portkey GatewayGitHubstargateway, guardrails, routingAI gateway with routing and guardrails for multi-model production traffic.
CAI (Cybersecurity AI)GitHubstarsecurity, governance, frameworkSecurity-focused agent framework for offensive/defensive AI workflows.
HigressGitHubstarai-gateway, mcp, governanceCNCF AI-native gateway for unified LLM API and MCP API management, including hosted remote MCP servers and gateway controls for agent tool access.
PlanoGitHubstarproxy, safety, data-planeAI-native proxy and data plane with orchestration, safety, and observability.
OpenAI Realtime AgentsGitHubstarrealtime, orchestration, controlAdvanced agentic realtime patterns with structured control and interaction loops.
OpenAI CS Agents DemoGitHubstardemo, handoffs, governanceCustomer-service multi-agent demo highlighting handoffs and guardrail-like control points.
Agent Governance ToolkitGitHubstargovernance, policy, sandboxingRuntime governance toolkit that deterministically enforces agent policy, identity, sandboxing, and audit controls before actions execute.
AgentGatewayGitHubstargateway, mcp, proxyAgentic proxy gateway for AI agents and MCP server ecosystems.
ContextForgeGitHubstargateway, governance, observabilityRegistry and proxy layer that unifies MCP, A2A, and REST/gRPC endpoints with centralized governance and observability.
ArchestraGitHubstarenterprise, guardrails, governanceEnterprise AI platform with guardrails, MCP registry, and orchestration services.
nonoGitHubstarpolicy, sandbox, auditCapability-based, policy-governed runtime that narrows agent access to explicit filesystem, network, credential, operation, sandbox, and audit capabilities.
TracecatGitHubstarsecurity, automation, policyAI automation platform for security teams with policy and workflow controls.
Snyk Agent ScanGitHubstarsecurity-scanner, mcp, skillsSecurity scanner that inventories local agent harnesses, MCP servers, and skills, then checks for prompt injection, tool poisoning, malware payloads, and sensitive-data risks.
Agent VaultGitHubstarcredentials, egress-policy, proxyCredential proxy and vault that brokers agent API access without exposing real secrets, with egress filtering and request logging.
ClawManagerGitHubstarcontrol-plane, governance, runtimesKubernetes-native control plane for governing agent runtimes, AI gateway access, and reusable skills across multiple agent backends.
HaftGitHubstargovernance, decisions, mcpDecision-governance harness that records falsifiable contracts, evidence, and commissions before agents execute.
ClawKeeperGitHubstarsafety, runtime-monitoring, governanceSafety framework for autonomous agents combining skill-based policies, runtime plugins, and watcher-based governance.
DefenseClawGitHubstargovernance, runtime-guardrails, auditSecurity governance stack for OpenClaw and agentic runtimes, combining admission control, runtime inspection, CodeGuard checks, sandbox policy, audit logs, and observability export.
CordumGitHubstargovernance, approvals, auditAgent control plane for pre-execution policies, approval gates, safety-kernel verdicts, live monitoring, audit trails, and Claude Code compliance firewall hooks.
SponsioGitHubstarcontracts, runtime-safety, guardrailsRuntime enforcement layer that checks every agent action against deterministic contracts before execution.
DashClawGitHubstarapprovals, policy, auditGovernance layer that intercepts risky agent actions, enforces policy, routes approvals, and records audit-ready decision trails.
TandemGitHubstarruntime-authority, approvals, auditGoverned runtime authority layer for agents with scoped execution, tool visibility, permissioned memory, approval gates, and audit trails.

Reference Harness Implementations

ProjectLinkStarsTagsSummary
OpenClawGitHubstargateway, channels, sandboxingLocal-first personal assistant harness with a gateway control plane for sessions, channels, tools, events, skills, and sandboxed non-main agents.
Hermes AgentGitHubstarmemory, skills, subagentsSelf-improving agent runtime with memory, skill creation, subagents, scheduled automations, and pluggable terminal backends.
DeepSeek HarnessGitHubstarplugin-architecture, session-log, sandboxingDeepSeek's official plugin-first agent harness with replaceable agent loops, model adapters, tool registries, durable session logs, sandbox and approval policies, and web or headless profiles.
OpenCodeGitHubstarterminal, coding-agent, subagentsOpen-source coding agent with built-in plan/build roles, subagents, LSP support, and a client-server runtime.
Claw CodeGitHubstarrust, cli, sessionsPublic Rust implementation of the claw CLI agent harness with auth, sessions, parity checks, container workflows, and terminal execution guidance.
Claude CodeGitHubstarterminal, coding-agent, git-workflowsOfficial terminal coding agent that understands codebases and executes editing, debugging, and Git workflows through natural language.
Codex CLIGitHubstarterminal, coding-agent, local-executionTerminal-native coding agent that runs locally and exposes practical agent workflows for software tasks.
Browser UseGitHubstarbrowser-agent, automation, benchmarksBrowser-agent framework that exposes websites to LLMs through browser state, tools, cloud browsers, and benchmarked task runs.
Gemini CLIGitHubstarterminal, coding-agent, mcpOpen-source terminal agent with built-in tools, MCP support, checkpointing, and sandboxing controls.
piGitHubstarcoding-agent, runtime, monorepoAgent harness monorepo combining a coding-agent CLI, shared runtime, and multi-provider LLM stack.
OpenHandsGitHubstarcoding-agent, software-engineering, repoOpen-source AI software engineer focused on repo-level coding task execution.
LobeHubGitHubstaroperator, multi-agent, schedulingChief-agent-operator platform for scheduling, running, and reporting on multi-agent workstreams.
PaperclipGitHubstarmanaged-agents, control-plane, governanceManaged-agent control plane with org charts, ticketing, budgets, heartbeats, and audit trails for coordinating agent teams.
learn-claude-codeGitHubstartutorial, harness, claude-codeHands-on harness tutorial for building Claude Code-like systems from scratch.
Open InterpreterGitHubstarterminal, coding-agent, harness-emulationTerminal coding agent with selectable harness emulation, native sandboxing, provider switching, skills, hooks, permissions, MCP, and AGENTS.md support.
ClineGitHubstarcoding-agent, mcp, checkpointsOpen-source coding agent spanning IDE, terminal, SDK, and kanban surfaces with shared approvals, MCP, checkpoints, and agent teams.
OpenManusGitHubstargeneral-agent, autonomy, workflowsOpen foundation for broad autonomous agent workflows with coding-heavy use cases.
gooseGitHubstardesktop, cli, mcpLinux Foundation AAIF open-source AI agent spanning desktop, CLI, and embeddable API with provider adapters, ACP access, and MCP extensions.
CLI-AnythingGitHubstarcli, tool-use, automationCLI agent system that unifies command-line tool usage in agent loops.
aiderGitHubstarterminal, repo-map, testingTerminal coding assistant with repo mapping, git-aware edits, and built-in lint/test feedback loops.
MulticaGitHubstarmanaged-agents, coding-agent, runtimesManaged-agents platform that assigns issues to coding agents, routes execution through runtimes, and compounds reusable skills.
nanobotGitHubstarruntime, memory, multi-channelUltra-lightweight agent runtime with WebUI, chat channels, tools, memory, MCP, model routing, deployment, and long-running goal support.
CowAgentGitHubstarreference, skills, multi-channelReference agent harness implementation with planning, memory, knowledge, skills, tools, MCP integration, schedulers, browser automation, and multi-channel delivery.
CodeWhaleGitHubstarcoding-agent, approval-gates, runtime-policyLocal-first coding-agent harness with explicit authority ordering, evidence loops, approval-gated tools, side-git snapshots, rollback, subagents, and runtime APIs.
Claude Code Plugins: Orchestration and AutomationGitHubstarclaude-code, plugins, orchestrationProduction-ready Claude Code plugin marketplace bundling agents, skills, tools, and multi-agent workflow orchestrators.
oh-my-claudecodeGitHubstarclaude-code, multi-agent, worktreesTeam-first orchestration layer for Claude Code with staged multi-agent execution, worktree-aware setup, and persistent session artifacts.
Agent TARSGitHubstarcomputer-use, browser-agent, mcpMultimodal computer and browser agent stack with CLI/Web UI, hybrid GUI/DOM browser control, MCP tools, event streams, and sandbox support.
DeepSeek-ReasonixGitHubstardeepseek, prefix-cache, coding-agentDeepSeek-native coding-agent harness optimized for prefix-cache stability, with separate planner and executor sessions, cache-aware compaction, MCP plugins, permissions, workspace sandboxing, and per-turn checkpoints.
oh-my-codexGitHubstarcodex, workflow, worktreesWorkflow layer for OpenAI Codex CLI with stronger session startup, standard planning-to-completion flows, durable state, skills, hooks, and worktree launches.
ZeroClawGitHubstarruntime, approval-gates, sandboxingSingle-binary agent runtime with providers, channels, tools, memory, SOPs, approval gates, sandboxing, ACP, and tool receipts.
NanoClawGitHubstarcontainers, claude-sdk, schedulingContainer-isolated Claude agent harness with channel routing, scheduled jobs, per-group memory, and small-codebase customization.
oh-my-piGitHubstarterminal, lsp, subagentsTerminal AI coding agent with edit safety, LSP integration, and subagent support.
Vibe KanbanGitHubstarcoding-agent, workspaces, reviewKanban control plane for planning, running, reviewing, and merging work from coding agents in isolated workspaces.
CrushGitHubstarterminal, coding-agent, mcpTerminal coding agent that wires project tools, code, workflows, sessions, LSP context, model switching, and MCP extensions into an LLM runtime.
Qwen CodeGitHubstarterminal, coding-agent, cliTerminal-native open-source coding agent tuned for practical dev loops.
Kilo CodeGitHubstarcoding-agent, ide, autonomousOpen-source coding agent spanning VS Code, JetBrains, CLI, cloud agents, reviews, custom modes, browser and terminal control, MCP marketplace, and autonomous CI runs.
cmuxGitHubstarmacos, workspace, browserNative macOS terminal and browser workspace for AI coding agents with notifications, split panes, and scriptable control.
Grok BuildGitHubstarcoding-agent, tui, acpOfficial Rust coding-agent harness and TUI with interactive, headless, and ACP modes plus MCP, skills, plugins, hooks, sandboxing, subagents, and long-running task support.
Compound EngineeringGitHubstarplugins, worktrees, reviewCross-agent engineering plugin that codifies brainstorming, planning, worktree execution, review, and knowledge compounding loops.
SuperClaude FrameworkGitHubstarconfig, personas, workflowConfiguration framework adding commands, personas, and method templates to coding agents.
OpenCodeReviewGitHubstarcode-review, deterministic-pipeline, verificationProduction-tested code-review agent harness combining deterministic file coverage and rule selection with isolated subagents, tool-driven context retrieval, comment verification, sessions, CI, and tracing.
SWE-agentGitHubstarswe, issue-fixing, toolingResearch-grade coding agent that resolves GitHub issues with explicit tooling loops.
DevikaGitHubstarassistant, planning, codingOpen-source coding assistant system for planning and implementing development tasks.
jcodeGitHubstarcoding-agent, terminal, rustRust coding-agent harness built for multi-session workflows, customization, memory, and terminal performance.
OpenFangGitHubstaragent-os, guardrails, rustRust agent operating system with autonomous capability packages, manifests, guardrails, tools, memory, sandboxing, audit trails, and channel adapters.
DeepCodeGitHubstarcoding-agent, verification-loop, durable-sessionsCoding-agent harness with durable sessions, goal-driven loops, evidence-based completion, permission and sandbox controls, reusable skills, parallel agents, context compaction, and headless automation.
OpenHarnessGitHubstartool-use, memory, multi-agentOpen agent harness implementation covering tool use, skills, memory, permissions, and multi-agent coordination.
PaseoGitHubstarcoding-agent, daemon, multi-deviceMulti-device coding-agent daemon and client stack for orchestrating local agents, parallel runs, and cross-provider workflows.
EigentGitHubstardesktop, cowork, productivityOpen-source desktop cowork agent for autonomous task execution and productivity.
AperantGitHubstarcoding-agent, parallel, memoryAutonomous multi-agent coding framework with parallel execution, isolated workspaces, QA loops, and persistent memory.
Learn Harness EngineeringGitHubstartutorial, verification, workflowsProject-based harness engineering course with runnable templates and staged implementations for instructions, state, verification, observability, automated loops, graph workflows, rollback, and human approval.
SupersetGitHubstarworktrees, desktop, parallelWorktree-based desktop orchestrator for running and reviewing parallel CLI coding agents from one workspace.
IronClawGitHubstarsecurity, wasm, routinesSecurity-first personal agent harness with WASM sandboxing, routines, tool plugins, and persistent memory.
Agent SGitHubstarcomputer-use, gui-agent, evaluationOpen-source computer-use agent framework with grounding models, reflection, local code execution option, and OSWorld-style evaluation support.
GitHub Copilot CLIGitHubstarterminal, coding-agent, mcpOfficial terminal coding agent built on GitHub's Copilot harness with MCP extensibility, approval controls, and GitHub-native context.
holaOSGitHubstarlong-horizon, desktop, durable-stateDesktop-first long-horizon agent environment with runtime, memory, tools, apps, and durable state.
ML InternGitHubstarml-agent, sandbox-tools, tracesHugging Face ML engineering agent with local and sandbox tool runtimes, private Space sandboxes, trace logging, and CLI execution.
Agent OrchestratorGitHubstarworktrees, parallel, dashboardWorktree-based orchestration layer for parallel coding agents with autonomous CI and review feedback handling.
Open SWEGitHubstarasync, coding-agent, sweAsynchronous open-source coding agent focused on software issue workflows.
GitHub Copilot SDKGitHubstarsdk, agent-runtime, json-rpcOfficial multi-language SDK for embedding the Copilot CLI agent runtime with sessions, planning, tools, file edits, permissions, custom agents, skills, MCP, hooks, and JSON-RPC lifecycle management.
HarnessGitHubstarclaude-code, meta-factory, agent-teamsClaude Code meta-factory that generates domain-specific agent teams, skills, orchestration patterns, and validation steps from a project description.
OSAURUSGitHubstarmacos, local-first, memoryNative macOS harness for autonomous coding agents with persistent memory.
mini-swe-agentGitHubstarminimal, swe, coding-agentMinimal coding agent implementation with strong benchmark competitiveness.
YuxiGitHubstarmulti-tenant, knowledge-graph, mcpMulti-tenant agent harness platform combining RAG, knowledge graphs, LangGraph orchestration, skills, MCP, subagents, and sandbox tools.
WebwrightGitHubstarbrowser-agent, code-as-action, playwrightMinimal browser-agent harness that lets coding models solve web tasks by writing rerunnable Playwright scripts in a workspace.
Google Agents CLIGitHubstargoogle-cloud, lifecycle, skillsGoogle Cloud CLI and skill bundle that gives coding agents scaffold, evaluation, deployment, publishing, and observability workflows.
Munder DifflinGitHubstarmulti-agent, desktop, memoryDesktop multi-agent harness that wraps terminal-agent CLIs with hive mailboxes, shared memory, an orchestrator, approvals, worktrees, and telemetry.
1CodeGitHubstarcoding-agent, orchestration, worktreesDesktop-first coding-agent orchestrator with worktree isolation, background sandboxes, MCP tooling, and automation triggers.
AgentTeamsGitHubstarmulti-agent, human-in-the-loop, shared-stateCollaborative multi-agent OS with manager-worker coordination, shared state, human-in-the-loop oversight, pluggable runtimes, and Kubernetes deployment through Matrix rooms.
gptmeGitHubstarterminal, tools, mcpTerminal-native personal agent with local tools, shell and web access, provider-agnostic models, plugins, skills, MCP, guardrails, and autonomous loops.
AI-DLC WorkflowsGitHubstarworkflow-rules, quality-gates, steeringOfficial AWS workflow ruleset that steers coding agents through adaptive phases, quality gates, and IDE-specific context files.
TinyAGIGitHubstarteam-orchestration, autonomous, workflowsTeam-style agent orchestrator for one-person-company style autonomous workflows.
Open Claude CoworkGitHubstardesktop, ui, orchestrationDesktop coding cowork assistant that turns agent orchestration into GUI workflows.
Amazon Bedrock AgentCore SamplesGitHubstaraws, runtime, operationsOfficial sample suite for deploying and operating agents with runtime, gateway, memory, observability, evaluation, and policy layers.
MaestroGitHubstardesktop, worktrees, orchestrationDesktop command center for parallel coding agents with worktree isolation, queued tasks, auto-run playbooks, and reusable sessions.
Claude Code HarnessGitHubstarclaude-code, workflow, review-loopClaude Code development harness that turns agent work into a repeatable spec, plan, work, review, and release loop.
KimchiGitHubstarterminal, multi-model, project-stateTerminal coding-agent harness with multi-model orchestration roles, phase tracking, persistent project state, and subagent delegation.
Open CoworkGitHubstardesktop, sandbox, mcpDesktop agent app with VM-backed sandboxing, MCP connectors, GUI control, and built-in skill workflows.
Pilot ShellGitHubstarcodex, claude-code, quality-gatesClaude Code and Codex workflow harness with spec-driven planning, enforced TDD, persistent memory, quality gates, and reusable skills.
thClawsGitHubstarrust, workspace, skillsNative Rust agent workspace with shared agent loop, sessions, tools, skills, MCP, memory, hooks, and sandboxing.
mini-coding-agentGitHubstarcoding-agent, minimal, approvalsMinimal coding agent harness illustrating approvals, memory, bounded delegation, and durable transcripts.
FlockGitHubstardesktop, visual-workflow, sandboxDesktop multi-agent harness with visual workflows, local and sandboxed execution, tool approvals, VNC takeover, MCP, skills, and scheduled tasks.
MateClawGitHubstarself-hosted, approvals, channelsSelf-hosted multi-user agent harness with StateGraph reasoning, skills, MCP/ACP registry, approvals, audit trail, and channel adapters.
codex-autorunnerGitHubstarmeta-harness, tickets, long-runningMeta-harness that treats tickets as the control plane for long-running coding agents, with queue execution, hub UI, and chat notifications.
CUGAGitHubstarenterprise, policies, mcpEnterprise generalist agent harness with MCP/OpenAPI tools, policy controls, HITL gates, memory, skills, and hybrid API/browser execution.
CheetahClawsGitHubstarcoding-agent, python, mcpPython agent harness infrastructure for long-horizon, multi-model, tool-using coding assistants with MCP, skills, memory, approvals, checkpoints, and bridges.
DextoGitHubstarcoding-agent, sessions, mcpOpen agent harness for AI applications with YAML configs, stateful sessions, tool orchestration, memory, observability, permissions, and subagents.
OpenClaw.NETGitHubstardotnet, gateway, governanceNativeAOT-friendly .NET agent runtime and gateway with tools, memory, MCP, governance ledger, evidence bundles, and harness regression tests.
UtahGitHubstardurable-execution, event-driven, multi-channelInngest-powered durable agent harness with a think-act-observe loop, step-level retries, singleton concurrency, cancellation, and multi-channel adapters.

Essential Readings & Ecosystem Maps

ProjectLinkStarsTagsSummary
awesome-claude-codeGitHubstarawesome-list, claude-code, skillsCommunity collection of Claude Code skills, hooks, and orchestrator tooling.
Awesome Agent SkillsGitHubstarawesome-list, skills, cross-harnessCurated cross-harness map of official and community Agent Skills for Claude Code, Codex, Gemini CLI, Cursor, OpenCode, Copilot, and related hosts.
awesome-agentic-patternsGitHubstarawesome-list, patterns, designCatalog of reusable agentic design patterns and implementation motifs.
awesome-mcp-serversGitHubstarawesome-list, mcp, toolsCurated MCP server index for tool interoperability in agent systems.
awesome-harness-engineeringGitHubstarawesome-list, curation, harnessCurated list focused on harness engineering articles, benchmarks, and implementations.
12 Factor AgentsReference-reading, operations, principlesOperations-oriented principles for building maintainable production agents.
Agent Frameworks, Runtimes, and Harnesses, oh my!Reference-reading, langchain, architectureClear decomposition of framework vs runtime vs harness responsibilities.
An open-source spec for Codex orchestration: Symphony.Reference-reading, openai, orchestrationOpenAI's orchestration write-up on turning issue trackers into always-on control planes for coding agents.
Building a safe, effective sandbox to enable Codex on WindowsReference-reading, openai, sandboxingOpenAI's implementation report on enforcing filesystem and network boundaries for Codex on Windows with restricted tokens, SIDs, and a dedicated command runner.
Building Effective AI AgentsReference-reading, anthropic, agentsAnthropic's practical guidance on when to use workflows vs. autonomous agents and how to structure them.
Claude Agent SDK overviewReference-reading, claude, sdkOfficial Claude Agent SDK documentation for building agents on the Claude Code harness, including tools, MCP, permissions, hooks, subagents, and session flows.
Claude Code auto modeReference-reading, anthropic, permissionsAnthropic's write-up on classifier-backed approval delegation for safer high-autonomy coding-agent runs.
Code execution with MCPReference-reading, anthropic, mcpAnthropic's design notes on controlled code execution via MCP boundaries.
Demystifying Evals for AI AgentsReference-reading, evals, anthropicMethodology for designing robust agent evals in non-deterministic trajectories.
Effective context engineering for AI agentsReference-reading, context, anthropicGuidance on context-window budgeting and working-state management for agents.
Effective harnesses for long-running agentsReference-reading, long-running, anthropicPractical guide to maintaining state, resumability, and reliability over long agent runs.
Evaluating Deep Agents: Our LearningsReference-reading, langchain, evaluationLangChain's practical lessons on evaluating stateful and long-horizon agents.
From model to agent: Equipping the Responses API with a computer environmentReference-reading, openai, executionOpenAI's engineering guide to agent execution with the Responses API, hosted shell containers, orchestration loops, compaction, skills, persistent files, and restricted networking.
Harness design for long-running application developmentReference-reading, app-dev, anthropicFollow-up article on improving long-running app generation through harness structure.
Harness Engineering (Martin Fowler)Reference-reading, architecture, fowlerArchitectural perspective on harness engineering and entropy control.
Harness engineering (OpenAI)Reference-reading, methodology, openaiField report on building reliable agent-first software via harness constraints and verification.
How GPT-5.6 fuses frontier intelligence with frontier efficiencyReference-reading, openai, efficiencyOpenAI's engineering account of the Rust agentic harness behind Codex and ChatGPT Work, covering deferred tool discovery, bounded outputs, deterministic tool ordering, and prompt-cache preservation.
How we built our multi-agent research systemReference-reading, anthropic, multi-agentAnthropic architecture write-up on role separation and coordination in multi-agent systems.
How we contain Claude across productsReference-reading, anthropic, containmentAnthropic's cross-product containment architecture for claude.ai, Claude Code, and Cowork, spanning sandboxes, VMs, filesystem boundaries, egress controls, and approval fatigue.
Improving Deep Agents with harness engineeringReference-reading, langchain, harnessEvidence that harness improvements alone can move benchmark performance.
Introducing Agent Executor, Google’s distributed Agent RuntimeReference-reading, google, runtimeGoogle's distributed agent runtime design for durable execution, event-log snapshots, secure isolation, session consistency, reconnection, and trajectory branching.
Making Claude Code more secure and autonomous with sandboxingReference-reading, anthropic, sandboxingHow Anthropic uses sandbox boundaries to raise agent autonomy without giving up security controls.
Quantifying infrastructure noise in agentic coding evalsReference-reading, anthropic, evaluationAnalysis of how infrastructure choices impact coding-agent benchmark outcomes.
Running Codex safely at OpenAIReference-reading, openai, governanceOpenAI's operational blueprint for governing Codex with sandbox and approval policies, network controls, managed identity, configuration, telemetry, and audit trails.
Scaling Managed Agents: Decoupling the brain from the handsReference-reading, anthropic, architectureAnthropic's meta-harness architecture for decoupling session logs, harness loops, and sandboxes in long-horizon agents.
Skill Issue: Harness Engineering for Coding AgentsReference-reading, humanlayer, coding-agentsPractical breakdown of why coding-agent quality depends heavily on harness setup.
Testing Agent Skills Systematically with EvalsReference-reading, openai, evalsOpenAI Developers guide for turning agent traces into repeatable skill evaluations.
The Anatomy of an Agent HarnessReference-reading, architecture, langchainConceptual decomposition of agent harness components and their responsibilities.
The next evolution of the Agents SDKReference-reading, openai, sdkOpenAI's product and engineering post on model-native agent harnesses, native sandbox execution, manifests, memory, and filesystem and shell tools.
Unlocking the Codex harness: how we built the App ServerReference-reading, openai, architectureOpenAI's architecture deep dive into sharing the Codex harness across clients through App Server, thread persistence, JSON-RPC, sandboxed tools, MCP, skills, and reconnectable sessions.
Unrolling the Codex agent loopReference-reading, openai, architectureOpenAI engineering deep dive into the Codex harness loop, prompt growth, tool-call replay, and stateless execution tradeoffs.
What We Learned Building Cloud AgentsReference-reading, cognition, cloud-agentsCognition's field report on secure cloud-agent infrastructure, VM isolation, full-state snapshots, orchestration, governance, integrations, and enterprise adoption.
Writing effective tools for AI agentsReference-reading, anthropic, toolsBest practices for tool interface design so agents call tools safely and reliably.
Your Agent Needs a Harness, Not a FrameworkReference-reading, inngest, reliabilityArgument for reliability-first infrastructure around agents instead of framework-only thinking.

Maintenance Notes

  • Source of truth: data/projects.yaml
  • Regenerate README files: python3 scripts/render_readme.py
  • Verify catalog and links: python3 scripts/verify_catalog.py

Citation

@misc{li2026agentharness,
  title={Agent Harness Engineering: A Survey},
  author={Li, Junjie and Xiao, Xi and Zhang, Yunbei and Liu, Chen and Zhao, Lin and Liao, Xiaoying and Ji, Yingrui and Wang, Janet and Gu, Jianyang and Ge, Yingqiang and Xu, Weijie and Fang, Xi and Xu, Xiang and Zhao, Tianchen and Kim, Youngeun and Wang, Tianyang and Hamm, Jihun and Krishnaswamy, Smita and Huan, Jun and Reddy, Chandan},
  url={https://openreview.net/pdf?id=eONq7FdiHa},
  year={2026}
}