Kandev System Architecture
July 26, 2026 · View on GitHub
Note: This document describes both the current implementation and the planned architecture. Sections marked with ✅ are implemented, while sections marked with 📋 are planned for future development.
Current Implementation Status
The current implementation uses a unified binary architecture instead of separate microservices:
- ✅ Single Go binary (
cmd/kandev/main.go) running all services - ✅ SQLite database for persistence (instead of PostgreSQL)
- ✅ In-memory event bus (instead of NATS)
- ✅ Docker-based agent execution with ACP streaming
- ✅ Native ACP (Agent Communication Protocol) - JSON-RPC 2.0 over stdin/stdout
- ✅ Permission request handling - Auto-approval of workspace indexing and tool permissions
- ✅ Session resumption for multi-turn agent conversations
- ✅ WebSocket-first API for all operations (replacing REST)
Overview
Kandev uses an event-driven architecture with Agent Communication Protocol (ACP) for real-time streaming between AI agents and the backend. The system is designed for deployment on client machines (local workstations) with future cloud scalability.
Current Architecture (Unified Binary)
flowchart TB
subgraph Binary["Kandev Binary (Port 38429)"]
Task["Task Service"]
Agent["Agent Manager"]
Orch["Orchestrator Service"]
Task & Agent & Orch --> EventBus["In-Memory Event Bus"]
EventBus --> DB["SQLite Database (kandev.db)"]
end
Binary --> Docker["Docker Engine"]
subgraph Container["Augment Agent Container"]
ACP["ACP via stdout"]
end
Docker --> Container
3. Orchestrator Service
Port: 8082 Purpose: Automated task orchestration, agent coordination, and real-time ACP streaming
Responsibilities:
- Subscribe to task state change events
- Determine when to launch AI agents
- Manage task queue with priority
- Coordinate with Agent Manager for container launches
- Aggregate ACP messages from NATS event bus
- Provide WebSocket endpoints for real-time ACP streaming to frontend
- Expose comprehensive REST API for task execution control
- Implement retry logic for failed tasks
- Monitor agent execution
Dependencies:
- NATS (event subscription + ACP message aggregation)
- Agent Manager (gRPC calls)
- PostgreSQL (read task details, store ACP messages)
Events Subscribed:
task.state_changedagent.completedagent.failedacp.message.*- All ACP messages from agents
WebSocket Actions:
orchestrator.subscribe- Subscribe to real-time ACP streaming for a taskorchestrator.unsubscribe- Unsubscribe from ACP streamingorchestrator.start- Start agent execution for a taskorchestrator.stop- Stop agent execution for a taskorchestrator.status- Get execution status for a task
Orchestration Logic:
When task.state_changed to "TODO" with agent_type set:
1. Add task to priority queue
2. Check available resources
3. Request Agent Manager to launch container
4. Update task state to "IN_PROGRESS"
5. Subscribe to ACP messages for this task
6. Stream ACP messages to connected WebSocket clients
When acp.message received:
1. Store message in database
2. Broadcast to WebSocket clients subscribed to this task
3. Update task progress/status
When agent.completed:
1. Update task state to "COMPLETED"
2. Record completion time
3. Send final ACP result message
4. Clean up resources
When agent.failed:
1. Check retry count
2. If retries available: re-queue task
3. Else: Update task state to "FAILED"
4. Send ACP error message
4. Agent Manager Service
Port: 8083 Purpose: Agent-environment lifecycle management with ACP streaming and scoped credential delivery
Responsibilities:
- Launch Docker containers for AI agents
- Resolve profile credentials and workspace-scoped GitHub credential leases for each launch
- Capture and parse ACP messages from container stdout
- Publish ACP messages to NATS event bus
- Monitor container health and status
- Stream and store container logs
- Manage agent type registry
- Resource allocation and limits
- Container cleanup
Dependencies:
- Docker Engine (container operations)
- PostgreSQL (agent_instances, agent_logs, agent_types tables)
- NATS (event publishing + ACP message publishing)
- Host filesystem and configured executor providers
Events Published:
agent.startedagent.runningagent.completedagent.failedagent.stoppedacp.message.{task_id}- Real-time ACP messages from agents
Container Launch Flow with ACP:
1. Receive launch request (task_id, agent_type)
2. Lookup agent_type configuration
3. Create agent_instance record (status: PENDING)
4. Pull Docker image if needed
5. **Resolve explicit profile credentials and repository-scoped GitHub broker leases**
6. **Checkout repository on host (if repository_url provided)**
7. Prepare volume mounts (workspace, credentials)
8. Create container with resource limits and mounts
9. Start container
10. Update agent_instance (status: RUNNING, container_id)
11. **Attach to container stdout for ACP message capture**
12. **Parse ACP messages from stdout**
13. **Publish ACP messages to NATS (acp.message.{task_id})**
14. Store ACP messages in database
15. Publish agent.started event
16. Monitor container until completion
17. Collect exit code and final logs
18. Update agent_instance (status: COMPLETED/FAILED)
19. Publish completion event
20. Cleanup container and workspace
Credential Delivery Strategy:
Profile Configuration:
- Literal or secret-referenced environment variables are explicit, unmanaged grants
- Executor-specific credential files are copied or mounted only when selected
Managed GitHub Automation:
- One workspace automation identity is resolved for the task
- A task/repository/generation-bound lease is supplied instead of an ambient host token
- Git uses agentctl's credential helper to redeem the matching repository lease on demand
- gh uses a PATH-prepended shim and the primary repository lease
- App private keys and personal GitHub tokens never enter the executor
Workspace Mount (Read-Write):
- /tmp/kandev/workspaces/{task_id} → /workspace
GitHub App installation tokens redeemed through the broker are minted for the requested
repository. PAT and named gh CLI credentials remain bearer tokens with their provider-granted
scope after redemption; lease and helper matching prevents accidental redemption for another
repository but does not cryptographically narrow a PAT or CLI token. Agent subprocesses are
therefore trusted with the workspace automation grant. Only migration-only legacy_shared
connections may resolve ambient backend gh, GITHUB_TOKEN, or GH_TOKEN authentication.
GitHub App root credentials live in a deployment catalog of registrations, but no registration is globally active. A workspace App connection names both one registration and one verified installation. Related workspaces may deliberately reuse a registration, sharing its bot identity, root credentials, permission policy, and root-credential revocation boundary; separate registrations provide independent work, personal, or organization trust boundaries. Installation tokens, repository grants, workspace generations, personal OAuth tokens, and broker leases remain workspace-scoped even when a registration is shared.
Manifest creation, existing-App import, installation, and personal authorization start from the
workspace GitHub settings flow. Runtime clients and caches are keyed by registration ID and
credential generation. Public callbacks and webhooks use
/api/v1/github/app/registrations/{registrationId}/...; the route ID selects one candidate App,
while single-use state or webhook HMAC still provides authorization. App credentials are persisted
as versioned encrypted bundles and are never sourced from backend startup configuration.
Data Flow Examples
Example 1: User Creates a Task with AI Agent (with Real-time ACP Streaming)
1. User → API Gateway: WebSocket message: task.create
{
"action": "task.create",
"payload": {
"boardId": "{id}",
"title": "Fix login bug",
"description": "Users can't login with email",
"agent_type": "auggie-cli",
"repository_url": "https://github.com/user/repo",
"branch": "main"
}
}
2. API Gateway → Task Service: Forward request
3. Task Service:
- Validate request
- Create task in database (state: TODO)
- Publish event to NATS: task.created
4. Orchestrator (subscribed to task.created):
- Receive event
- Check if agent_type is set
- Add task to priority queue
- Process queue
5. Orchestrator → Agent Manager (gRPC):
LaunchAgent(task_id, agent_type="auggie-cli")
6. Agent Manager:
- Create agent_instance record
- Pull docker image "augmentcode/auggie-cli:latest"
- **Resolve profile credentials and GitHub broker leases**
- **Clone repository to /tmp/kandev/workspaces/{task_id}**
- Create container with volume mounts:
* /tmp/kandev/workspaces/{task_id} → /workspace
* task-scoped executor credentials and repository leases
- Start container
- **Attach to container stdout for ACP capture**
- Publish agent.started event
7. Orchestrator (subscribed to agent.started):
- Update task state to IN_PROGRESS
- **Subscribe to acp.message.{task_id}**
8. Frontend → API Gateway: WebSocket connection
WS /api/v1/orchestrator/tasks/{task_id}/stream
9. API Gateway → Orchestrator: Proxy WebSocket connection
10. Agent Container → stdout:
**Writes ACP messages in JSON format:**
{"type":"progress","timestamp":"...","agent_id":"...","task_id":"...","data":{"progress":10,"message":"Analyzing codebase..."}}
11. Agent Manager:
- **Captures ACP message from container stdout**
- **Parses JSON ACP message**
- **Publishes to NATS: acp.message.{task_id}**
- Stores ACP message in database
12. Orchestrator (subscribed to acp.message.{task_id}):
- **Receives ACP message from NATS**
- **Broadcasts to all WebSocket clients for this task**
13. Frontend (via WebSocket):
- **Receives real-time ACP message**
- **Updates UI with progress: "Analyzing codebase... 10%"**
14. [Steps 10-13 repeat for each ACP message from agent]
15. Agent Container completes:
- Writes final ACP result message
- Exits with code 0
16. Agent Manager:
- Captures final ACP message
- Publishes to NATS
- Collects exit code
- Publishes agent.completed event
17. Orchestrator (subscribed to agent.completed):
- Update task state to COMPLETED
- Record completion time
- **Send final ACP result to WebSocket clients**
- Close ACP subscription
18. Frontend:
- **Receives completion notification**
- **Displays final results**
- Closes WebSocket connection
Example 2: Manual Task State Change
1. User → API Gateway: WebSocket message: task.state
{
"action": "task.state",
"payload": {
"taskId": "{id}",
"state": "IN_PROGRESS"
}
}
2. API Gateway → Task Service: Forward request
3. Task Service:
- Validate state transition
- Update task state
- Create task_event record
- Publish task.state_changed event
4. Orchestrator (subscribed to task.state_changed):
- Check if agent should be launched
- If no agent_type set: ignore
- If agent_type set: add to queue
Agent Communication Protocol (ACP)
Protocol Overview
Kandev uses the Agent Communication Protocol (ACP) - an open protocol for communication between AI agents and client applications. The implementation follows the official ACP specification using JSON-RPC 2.0 over stdin/stdout.
Reference: https://agentclientprotocol.com
Key Characteristics:
- JSON-RPC 2.0: Standard request/response/notification format
- Bidirectional: Client sends requests, agent sends responses/notifications/requests
- Session-based: All communication happens within a session context
- Permission model: Agent requests permissions before tool execution
Current Implementation ✅
The backend implements full ACP support with:
-
Session Management (
apps/backend/internal/agent/acp/session.go)- Create sessions with
session/new - Send prompts with
session/prompt - Resume sessions with stored session IDs
- Create sessions with
-
JSON-RPC Client (
apps/backend/pkg/acp/jsonrpc/client.go)- Handles responses to client requests
- Handles notifications from agent (
session/update) - Handles requests from agent (
session/request_permission)
-
Permission Handling
- Auto-approves workspace indexing permissions
- Auto-approves tool execution permissions (selects first "allow" option)
ACP Message Types
1. Client → Agent Requests
// session/new - Create a new session
{
"jsonrpc": "2.0",
"id": 1,
"method": "session/new",
"params": {
"cwd": "/workspace",
"mcpServers": []
}
}
// session/prompt - Send a prompt to the agent
{
"jsonrpc": "2.0",
"id": 2,
"method": "session/prompt",
"params": {
"sessionId": "uuid",
"content": [{"type": "text", "text": "Analyze this code"}]
}
}
2. Agent → Client Responses
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"sessionId": "uuid",
"status": "ok"
}
}
3. Agent → Client Notifications
// session/update - Progress updates
{
"jsonrpc": "2.0",
"method": "session/update",
"params": {
"sessionId": "uuid",
"content": [{"type": "text", "text": "Analyzing..."}],
"stopReason": null
}
}
// session/update - Completion
{
"jsonrpc": "2.0",
"method": "session/update",
"params": {
"sessionId": "uuid",
"content": [{"type": "text", "text": "Done!"}],
"stopReason": "end_turn"
}
}
4. Agent → Client Requests (Permissions)
// session/request_permission - Agent requests permission
{
"jsonrpc": "2.0",
"id": 5,
"method": "session/request_permission",
"params": {
"sessionId": "uuid",
"toolCall": {
"toolCallId": "workspace-indexing-permission",
"title": "Workspace Indexing Permission"
},
"options": [
{"optionId": "enable", "name": "Enable indexing", "kind": "allow_always"},
{"optionId": "disable", "name": "Disable indexing", "kind": "reject_always"}
]
}
}
// Client responds with selected option
{
"jsonrpc": "2.0",
"id": 5,
"result": {
"outcome": {
"outcome": "selected",
"optionId": "enable"
}
}
}
ACP Flow Through System (Current Implementation)
flowchart TB
subgraph Backend["Kandev Backend"]
Orch["Orchestrator<br/>Execute(task)"]
ACP["ACP Session Manager<br/>Create / Prompt / Permission"]
RPC["JSON-RPC Client<br/>Send requests / Handle responses"]
Orch --> ACP --> RPC
end
RPC <-->|"Docker attach (stdin/stdout)"| Agent
subgraph Agent["Agent Container"]
CLI["Auggie CLI (--acp)"]
stdin["stdin ← JSON-RPC requests"]
stdout["stdout → responses/notifications"]
CLI --- stdin
CLI --- stdout
end
Workspace["/workspace<br/>Mounted from host"] --> Agent
Initial Agent Types
1. Auggie CLI Agent
- Image:
augmentcode/auggie-cli:latest - Capabilities: Code analysis, generation, debugging, refactoring
- ACP Support: Native
- Resources: 2 CPU, 2GB RAM
2. Gemini Agent
- Image:
kandev/gemini-agent:latest - Capabilities: Code review, documentation, testing
- ACP Support: Wrapper implementation
- Resources: 1.5 CPU, 1.5GB RAM
Database Schema Summary
Tables:
users- User accountsboards- Kanban boardstasks- Tasks on boardstask_events- Audit log of task changesagent_types- Registry of available agent typesagent_instances- Running/completed agent containersagent_logs- Logs from agent execution
Key Relationships:
- Board → Tasks (1:N)
- Task → Agent Instances (1:N)
- Agent Instance → Agent Logs (1:N)
- User → Boards (1:N)
- User → Tasks (created_by)
Technology Stack
Backend Services:
- Language: Go 1.21+
- Web Framework: Gin
- Database: PostgreSQL 15+
- Event Bus: NATS
- Container Runtime: Docker
Key Libraries:
github.com/docker/docker- Docker SDKgithub.com/gin-gonic/gin- HTTP frameworkgithub.com/gorilla/websocket- WebSocket supportgithub.com/jackc/pgx/v5- PostgreSQL drivergithub.com/nats-io/nats.go- NATS clientgithub.com/golang-jwt/jwt/v5- JWT authgo.uber.org/zap- Structured logginggoogle.golang.org/grpc- gRPC communication
Deployment Architecture
Primary Target: Client Machines (Local Workstations)
The system is designed for deployment on client machines (developer workstations, local servers) with direct access to user credentials.
Architecture:
flowchart TB
subgraph Machine["User's Local Machine"]
subgraph Infra["Infrastructure (Docker)"]
PG["PostgreSQL"]
NATS["NATS Server"]
end
subgraph Services["Kandev Backend Services"]
Gateway["API Gateway"]
TaskSvc["Task Service"]
OrchSvc["Orchestrator"]
AgentMgr["Agent Manager"]
end
subgraph Docker["Docker Engine"]
Auggie["Auggie Agent"]
Gemini["Gemini Agent"]
More["..."]
end
Creds["Host Credentials (Read-Only)<br/>~/.ssh/ ~/.gitconfig ~/.git-credentials"]
end
Services --> Docker
Creds -.-> Docker
Deployment Methods:
-
Docker Compose (Development)
docker compose up -d -
Native Binaries (Production)
systemctl start kandev-gateway systemctl start kandev-task-service systemctl start kandev-orchestrator systemctl start kandev-agent-manager
System Requirements:
- OS: Linux (Ubuntu 20.04+) or macOS 12+
- CPU: 4+ cores (8+ recommended)
- RAM: 8GB minimum (16GB recommended)
- Disk: 50GB free space
- Docker: 20.10+
Future Deployment Options (Roadmap)
Cloud Deployment (Phase 7+):
- Deploy to AWS/GCP/Azure
- Managed PostgreSQL and NATS
- Cloud secret managers for credentials
- Horizontal scaling with load balancers
Kubernetes (Phase 8+):
- Helm charts for deployment
- StatefulSets for databases
- Deployments for services
- Horizontal Pod Autoscaling
- Ingress for external access
Multi-User SaaS (Phase 9+):
- Tenant isolation
- Centralized credential management
- Usage-based billing