Multi-source evidence architecture
September 10, 2026 ยท View on GitHub
Status: historical pre-implementation RFC (not a current operations or route reference) Version: 2 Last updated: 2026-07-23
This document preserves the design and migration rationale that preceded the
current implementation. For current behavior, use README.md, INSTALL.md, and
docs/reference.md: the event ledger is SQLite-backed by default, the HTML
dashboard and /raw route are retired, and interactive product surfaces are the
macOS app and agentacct tui over the local JSON API/store.
In particular, the RFC's original JSONL-only migration premise is superseded:
fresh and adopted stores use events.sqlite3 as the authoritative ledger
unless legacy flat-ledger mode is selected explicitly.
Decision
agentacct will present one evidence product while keeping a federated implementation. agentacct remains the evidence and reconciliation kernel. Mechanical client hooks, native client logs, CI, MCP, and provider records are independent sources which make bounded claims.
MCP remains a first-class, high-value semantic source. It is not removed or replaced. It also is not required for a client session, tool call, usage record, or machine result to exist. When MCP is unavailable, the product must still show the mechanically observed activity and explicitly mark work meaning as missing.
native logs host hooks CI
| | |
+-----------------+------------+
|
durable raw spool
|
versioned adapters
|
immutable EvidenceEnvelope v2
|
source policy + identity reconciliation
^
|
MCP / Work Event semantics
|
v
Work Graph | Evidence Matrix | Discrepancies | Cost & Outcome Basis
This architecture follows four rules:
- Raw evidence is append-only. A correction supersedes; it never overwrites.
- Every source is authoritative only for specific dimensions.
- A relationship is a claim with evidence, method, confidence, and conflicts.
- Missing, partial, estimated, and conflicting are product states, not errors to hide with query-time coalescing.
Product boundaries
agentacct owns
- stable local Task identity and explicit continuation/merge history;
- immutable evidence envelopes and source provenance;
- deterministic deduplication and replay;
- identity and link claims with conservative refusal;
- cross-source discrepancies;
- cost, usage, outcome, and artifact basis labels;
- local-first privacy defaults;
- advisory policy signals derived from reconciled evidence;
- a local control plane for Tasks and execution processes agentacct itself launches, including attempts, approvals, budgets, schedules, and registered workspaces.
Integrated systems keep owning
- model providers: provider usage and billing records;
- CI systems: objective check execution and status;
- host agents: lifecycle hooks and local transcripts/rollouts.
agentacct does not take over or mutate their processes. Its local control plane controls only agentacct-owned executions and projects their actions into the same Task/evidence model. It does not copy multi-tenant organization charts, RBAC, generic project management, tracing backends, code review workflows, or provider billing portals.
EvidenceEnvelope v2 contract
Every source normalizes into a vendor-neutral envelope.
Required identity and provenance:
evidence_idandidempotency_key;schema_version,source_system,source_type,source_instance_id;source_event_id,vendor_schema_version,adapter_version;event_kind,occurred_at,observed_at,ingested_at;raw_digest, optional localraw_ref, and envelopeintegrity_hash.
Required epistemic state:
measurement_basis;completeness:complete,partial, orunknown;- optional
truncation_reason; - usage and cost confidence;
- capture level and redaction profile;
- sensitive-field presence declaration.
Subject references are sparse and may include organization, project, principal,
work item, execution, client session, turn, tool call, trace, span, artifact,
commit, pull request, machine check, provider account, or billing record.
An identifier is never global by spelling alone. Product projections namespace
it by its issuing system, organization/project scope, and (except for native
client-issued ids) source instance. Equal raw ids from a provider or another
project remain separate until a validated ClaimedLink connects their evidence.
The envelope distinguishes facts from assertions:
observation: a source mechanically measured or saw something;claim: a source or agent asserted meaning, ownership, completion, or cost;derived: agentacct produced a reversible projection from immutable inputs.
An adapter may not label an event observation when it only received a user or
agent assertion. The store rejects invalid combinations rather than silently
upgrading their trust.
Link contract
Entity relationships are represented by ClaimedLink, not by an unqualified
foreign key:
from_entity -> relation -> to_entity
asserted_by_evidence_id
join_method
confidence
candidate_count
conflict_state
valid_from / valid_to
Hard invariants:
- no unique strong evidence means no
exactlink; - a conflict vetoes exact attribution;
- candidate count greater than one is ambiguous unless an independent exact key disambiguates it;
- a native usage log can prove session usage but not the business goal;
- MCP proves an agent-reported semantic statement, not provider billing.
Source authority by dimension
primary means the strongest available direct basis for that dimension.
corroborating can support or conflict but cannot replace a primary source.
claim-only is rendered as a claim even when the source is internally trusted.
| Source | Lifecycle/activity | Work meaning | Usage | Cost | Artifact | Outcome | Control |
|---|---|---|---|---|---|---|---|
| Native host hook | primary | unavailable | unavailable | unavailable | corroborating | corroborating when exit status is observed | unavailable |
| Native client log | primary | limited | primary client-reported | estimated or client-reported | corroborating | unavailable | unavailable |
| MCP Work Event | corroborating | primary claim | debug-only unless independently measured | debug-only | claim-only | claim-only unless linked to machine evidence | unavailable |
| CI check | corroborating | unavailable | unavailable | unavailable | corroborating | primary machine result | unavailable |
| Provider response | unavailable | unavailable | primary provider-reported | primary only when explicitly billed | unavailable | unavailable | unavailable |
| Provider invoice | unavailable | unavailable | aggregate/limited | primary billed | unavailable | unavailable | unavailable |
When two sources disagree about the same narrowly identified fact agentacct preserves both envelopes and creates a discrepancy. A whole session or run is not one immutable fact: sequential usage totals, changing lifecycle states, and repeated checks are compared only when they share a request/span/tool fact, an explicit measurement window, or the exact same event point. agentacct does not choose a winner outside the declared dimension policy.
Capture contract
Each adapter declares capabilities independently:
- session lifecycle;
- turns;
- tools;
- edits/artifacts;
- subagents;
- machine results;
- usage and cost;
- stable event identity;
- completeness and truncation reporting.
Hook execution follows a local-first hot-path contract:
- parse a bounded JSON payload;
- allowlist and redact metadata;
- compute deterministic source identity;
- append to the local durable spool;
- return without rebuilding product projections or making a network request.
Capture is fail-open for the host. A broken agentacct adapter must not block the developer's agent. The failure is recorded locally when possible and exposed by doctor/status commands.
The first canonical mechanical events are:
session_observed;turn_observed;tool_observed;machine_check_observed;session_link_observed.
Commands are machine checks only when an objective exit status exists and the command category is recognized as test, build, lint, or typecheck. Text such as "tests passed" is never parsed into a passed result without objective evidence.
Work Event contract
The semantic model is transport-neutral. Existing agentacct_* MCP tools remain
compatible and write Work Events through the same normalization boundary as
HTTP imports.
A Work Event may express:
- objective and work-item identity;
- section start, checkpoint, completion, blocker, or clean handoff;
- decision, rationale, and next step;
- asserted artifact or outcome meaning;
- exact client/session/turn/message join keys when the host exposes them.
The transport is part of provenance. An HTTP event does not become MCP evidence, and an MCP event does not become a mechanical observation. No generic writer can mint trusted local-usage, hook-observed, CI, or provider-billed provenance.
Storage and projection
Evidence v2 began as a shadow layer during migration. In the current build:
- the main event ledger is
events.sqlite3by default;events.jsonlis a legacy flat-ledger authority or an adopted-store transition backup only; - v2 raw envelopes use a dedicated append-only spool;
- trusted refreshable usage transitions use a second append-only spool,
evidence-v2/refreshable-usage.jsonl; - a SQLite projection indexes evidence identity, source, subject, status, and discrepancy candidates;
- replay is idempotent and does not re-add envelopes;
- acknowledgement changes processing state, not immutable envelope content;
- derived views can be deleted and rebuilt from the spool;
- disabling v2 stops shadow writes without modifying v1 behavior.
High-volume hook events must never be appended directly to the main event ledger. Retention and future downsampling apply to raw activity, never to provider billing or machine-outcome evidence without an explicit policy.
Trusted refreshable usage is a current-fact lane
Local usage importers periodically replace cumulative facts in the authoritative
event ledger. The event id and server write time are transport details of that replacement,
not logical usage identity. Treating either as source_event_id for every
refresh turns an unchanged cumulative fact into unbounded raw Evidence versions.
Evidence v2 therefore has one narrowly gated reconciliation lane for trusted refreshable usage. Only an allowlisted local usage importer may enter this lane, and only for cumulative current-state rows whose provenance and measurement basis pass source policy. A generic MCP, HTTP, or Work Event cannot opt itself into this authority by copying metadata.
The stable slot key contains all of:
- source namespace fingerprint (or an explicit unresolved sentinel that never coalesces with a resolved namespace);
- client;
- the client-gated normalized session id;
- importer-defined usage lane;
- representation class; and
- cumulative-update semantics.
Different clients, namespaces, sessions, lanes, representations, or additive
versus cumulative semantics are never folded together. The importer owns the
lane definition: when a client reports model-specific lanes, the model is part
of that lane; otherwise a provider/model change is a new revision of the same
slot rather than a new current total. This prevents both cross-lane merging and
double-counting a model switch. Product-facing synthetic run references retain
a readable bounded prefix plus a truncated 20-hex (80-bit) prefix of the slot
digest โ enough that truncation or sanitization of the readable part cannot
accidentally merge otherwise distinct current facts. The complete slot digest
is bound on the evidence envelope itself (source_event_id/raw_ref), not in
the run reference; auditors comparing slots must use the envelope digest, not
the run-reference suffix.
Each candidate also has a normalized truth digest. It includes the usage field-presence map and values, provider/model truth, usage confidence, client-reported cost and its basis, normalization/additivity state, hold/precedence state, and the exact supporting evidence references. It excludes the reminted event-ledger id, server-created/poll/ingest timestamps, raw location, scan order, and price-derived estimated cost. Repricing therefore does not manufacture a new client-usage fact, while changed client-reported cost or changed evidence lineage does.
Transitions are ordered only by a source-native revision/update watermark, never by the random event-ledger id, agentacct scan time, or server write time. A different revision with no comparable source order is retained as a conflict rather than guessed into sequence:
- the same truth digest is a physical no-op while it remains the current head;
- if that head was tombstoned, a later complete snapshot containing the fact appends one ordered resurrection transition rather than leaving it deleted;
- a strictly newer different digest appends one immutable revision and supersedes the prior head;
- an older candidate is stale and cannot roll the head back;
- divergent content at the same order becomes one stable, fail-visible conflict; retries do not append more copies; and
- a tombstone is an ordered transition, not an in-place deletion.
Slot comparison, transition append, and head projection occur under the same store lock, so concurrent identical refreshes create at most one transition and a delayed older refresh cannot win. A tombstone may be inferred only from a complete, successful, authoritative snapshot of the current trusted event ledger. A partial, limited, failed, or corrupt scan may project observed rows, but absence in that scan is never deletion evidence.
After every successful persisted watcher tick, agentacct reconciles the full current trusted usage slice in the authoritative event ledger, not only rows rewritten during that tick. This makes a prior fail-open Evidence shadow error self-healing on the next good tick. Broken ticks retain the previous head and cannot create tombstones.
The transition spool is separate from the generic spool.jsonl so a previous
binary never attempts to decode a new record kind. Each transition batch also
records the generic-spool byte fence observed under the shared lock, allowing a
compatible binary to replay both spools in the original cross-spool receipt
order. Generic EvidenceEnvelope behavior does
not change: its exact source identity still controls duplicate receipts and
same-identity/different-content conflicts, and it does not use refreshable-slot
coalescing.
For a clean rebuild, original Codex, Claude, and Hermes logs plus the trusted event ledger are the usage recovery inputs. Retained raw mechanical-capture inputs remain the recovery basis for their own evidence. Neither the inflated refresh history nor the SQLite projection is promoted to source truth merely because it is large; both old spools remain immutable archives until an owner accepts the rebuilt store.
Unified product views
The v2 product adds four evidence-specific views rather than cloning upstream dashboards:
- Work Graph: work item, execution, client session, trace, artifact, check, and billing relationships with link confidence.
- Evidence Matrix: which source asserts each dimension, including missing and partial coverage.
- Discrepancies: duplicate, conflict, missing-basis, and completion/outcome mismatches without automatic coalescing.
- Cost & Outcome Basis: displayed numbers and statuses grouped by client reported, estimated, provider reported, provider billed, and objective machine evidence.
The HTML dashboard described by the original shadow-period plan is retired. The evidence views remain available as bounded JSON projections and do not rebuild the entire event history per request.
Primary product projection
The normal user journey is deliberately smaller than the evidence model:
- Work leads with actionable reconciliation gaps, named work, recent agent activity, historical source coverage, and a compact usage snapshot.
- Usage provides the complete saved usage explorer.
- Sources and evidence inspection remain available through the TUI, CLI, and bounded JSON projections.
/sessions remains a stable JSON Work explorer. The current evidence routes are
namespaced under /evidence/*, including /evidence/work-graph,
/evidence/matrix, /evidence/discrepancies, and
/evidence/cost-outcome-basis. The former HTML /raw route is retired.
The primary projection follows additional honesty rules:
- source evidence proves historical coverage, never current source health;
connection_state=not_verifiedis neither connected nor disconnected;- agent/MCP work meaning remains a claim until linked mechanical evidence supports it;
- a trusted mechanical session observation may create an activity-only Task without MCP, but it carries no invented work title, token total, or cost;
- ambiguous usage is never allocated to make a work card look complete;
- opaque ids and normalized record metadata stay behind bounded evidence inspectors;
- large record APIs use bounded cursor pages rather than unbounded default responses.
Control bridge
Phase 6 is advisory by default. agentacct emits a structured ControlSignal
with its supporting evidence, basis, confidence, expiry, and recommended action.
Another controller decides whether to act.
Hard enforcement is not eligible unless:
- the evidence basis is provider-billed; or
- a user explicitly approved a conservative basis and threshold; and
- every supporting evidence id resolves in the same local store, passes integrity/target/conflict validation, and is bound to this signal; and
- the target controller owns the execution; and
- the signal is fresh, non-conflicting, and idempotent.
agentacct never pauses or cancels an external run merely because an agent claimed completion or an estimated price crossed a threshold.
Phase delivery and rollback
Phase 4 (read-only connectors) was removed in 0.9.1; the remaining phase numbers are unchanged.
Phase 0: contracts
Deliver this RFC, the privacy threat model, and golden scenarios. Rollback is documentation rejection; no runtime state changes.
Phase 1: evidence shadow kernel
Deliver envelope/link models, source policy, v1 adapter, spool, projection, dedupe, replay, and conflict preservation. Disable the feature to return to v1; delete only derived projection data, never raw evidence.
Phase 2: mechanical capture
Deliver Claude Code, Codex, and Cursor capability adapters and manifest rendering. Installation/config mutation is separate and explicit. Disable the capture adapter; current local import/MCP paths remain intact.
Phase 3: transport-neutral Work Events
Route semantic events through a common normalization service while preserving
all public agentacct_* names and v1 writes. Disable shadow normalization to use
only the existing MCP/API path.
Phase 5: evidence product
The original phase delivered the four API/UI projections. Its feature flag formerly returned to the then-current v1 HTML dashboard; that fallback is now retired. Derived views can still be rebuilt from immutable evidence.
Phase 6: advisory control and verification
Deliver advisory signals, refusal rules, replay/privacy/chaos/performance tests, and local dogfood. External mutation remains disabled until a later explicit authorization and product decision.
Phase 7: product projection and progressive disclosure
The original phase delivered the Work-first HTML dashboard, actionable attention cards, human-readable work summaries with closed evidence explainers, historical source coverage, an Advanced inspection hub, and bounded evidence-event pagination. The HTML surface and its legacy routes were later retired; immutable evidence and the bounded JSON projections remain available.
Phase 7.1: product home and action ownership
The original phase replaced the dashboard-style collection of attention, work, session, source,
and usage panels with one deduplicated Work feed. The product projection maps
work into In progress, Blocked, Open finding, Verified, Agent reported, or Activity observed. Only an explicit blocker or recorded
user-owned next action is a user-facing action; a failed machine check remains
an open finding until the agent resolves it and is not assigned to the user.
Missing semantic context, join keys, usage
attribution, and source-health proof are integration or diagnostic states; they
remain in diagnostic TUI/JSON views and never masquerade as a user task.
For project-scoped stores, Work excludes activity carrying an explicit other- project label while retaining older events whose project label is absent. Usage and evidence stores are unchanged. This phase is a presentation and scope-projection change, not a rewrite of ledger truth or immutable evidence.
Phase 7.2: mechanical session observation bridge
Project accepted Evidence v2 client-hook observations into bounded saved Session/Task activity. Exact client session and parent ids control folding; namespace conflicts fail closed. Observed agent/model labels and recognized machine checks may enrich the Task, while usage/cost lanes remain empty unless a supported local usage importer supplies them. MCP continues to provide named work meaning when it fires; the bridge does not synthesize it.
Go/no-go invariants
false exact = 0in the adversarial corpus;- replaying a fixture 100 times creates one immutable envelope per source event;
- overlapping cost claims are not summed;
null/unknown cost never becomes$0.00;- metadata-only capture contains no prompt, response, thought, tool argument, or tool result text;
- MCP absence never means that work or usage did not exist;
- an adapter cannot grant itself authority outside source policy;
- disabling all v2 flags leaves v1 behavior and storage unchanged;
- no integration introduces unknown or incompatible license material.
Deliberate non-goals
- replacing MCP with hooks;
- building a multi-tenant scheduler, RBAC, organization hierarchy, or generic project-management UI. agentacct may implement local approvals, registered workspaces, and bounded schedules for agentacct-owned attempts;
- storing full prompts, responses, thoughts, transcripts, or tool bodies by default;
- treating telemetry volume as product value;
- inferring exact business ROI from activity alone.