JavaScript Application Graph

July 16, 2026 ยท View on GitHub

JavaScript Application Graph v1 is REA's provider-neutral domain model for connecting shipped application bytes, JavaScript structure, Electron process boundaries, passive browser observations, and native add-ons. It is a durable, canonical data contract; it is not an extractor or a new CLI/MCP tool.

The implementation is pure and side-effect free. It validates caller-supplied facts, derives semantic identifiers, enforces graph integrity, and serializes verified graphs as RFC 8785 canonical JSON. It performs no filesystem or CDP I/O, does not evaluate JavaScript, and has no application-specific node kinds.

Layer boundary

The graph complements existing REA models instead of replacing them:

Existing modelResponsibilityApplication Graph use
Artifact graphContent-addressed containers, byte occurrences, and extraction boundariesSupplies artifact identities and exact paths
Web bundle analysisBounded AST, chunk, route, endpoint, and source-map observationsSupplies static nodes and relationships
Browser/Electron observationAuthorized passive CDP metadata and runtime capturesSupplies runtime instances and observed_as relationships
Native analysisProvider-qualified binary functions, symbols, and referencesSupplies native add-on/export observations
JavaScript Application GraphCross-layer entity identity, evidence-bearing relationships, and canonical commitmentsPreserves the combined application model

The shipped JavaScript artifact reconstruction service projects bounded directory/ASAR, package, bundle, source-map, BrowserWindow, preload, contextBridge, IPC, utility-process, and JavaScript-side native binding facts into this graph. The separate static/runtime reconciliation service adds capture-scoped targets, frames, scripts, workers, and inferred observed_as edges from already-produced Evidence. Native binary export verification remains a separate later layer. The domain graph itself never reaches outward to obtain any of those facts.

Versioned record

A stored graph has this top-level shape:

schema: "JavaScriptApplicationGraph"
schema_version: 1
graph_id: jag_<sha256 of normalized graph semantics>
root_node_ids: jag_node_<sha256>[]
nodes: application entities with one or more observations
edges: directed evidence-bearing relationships
coverage: complete, partial, unknown, or unavailable
limitations: explicit graph-wide constraints

Parsing rejects unknown fields, unsupported versions, non-canonical ordering, duplicate IDs, stale semantic commitments, missing roots, dangling endpoints, self-edges, and changed_from edges between unlike node kinds. Constructors normalize set-like arrays before deriving their IDs.

Entities and relationships

The v1 node vocabulary covers:

  • package, installer, and artifact;
  • ASAR entry;
  • Electron main, preload, renderer, and utility process code;
  • JavaScript asset, chunk, and module;
  • source map and original source module;
  • BrowserWindow, frame, and target;
  • contextBridge API, IPC channel, and IPC handler;
  • worker and service worker;
  • endpoint and storage;
  • native add-on and native export;
  • runtime script instance;
  • an explicit unknown kind when classification is not justified.

The directed relation vocabulary is contains, loads, imports, maps_to, exposes, sends, invokes, handles, calls, persists_to, observed_as, and changed_from. Provider names are deliberately absent from both vocabularies.

For example, a source-owned synthetic fixture proves this path without using a third-party application:

package contains app.asar
main asset contains BrowserWindow
BrowserWindow loads preload role
preload asset exposes contextBridge API
preload asset invokes uniquely matched main-process IPC handler
main-process IPC handler handles literal IPC channel
JavaScript asset imports requested native binding
resolved native add-on exposes requested binding (not a verified binary export)

The loads, exposes, invokes, handles, and imports relationships remain static inferences. Dynamic or ambiguous IPC is retained without a pairing edge. Runtime entities retain passive-CDP authority, while a successful static/runtime correspondence is a separately identified cross-layer inference.

Evidence and authority

Every observation and edge carries all of the following:

  • an authority, such as artifact bytes, AST analysis, static relationship inference, passive CDP runtime, cross-layer reconciliation, controlled replay, native analysis, historical reference, or user assertion;
  • an epistemic state: observed, inferred, unknown, or unavailable;
  • confidence independently bounded as exact, high, medium, low, or unknown;
  • an explicit content-addressed artifact reference or an unavailable reason;
  • an explicit artifact, source, URL, runtime, native, or package-export location, or an unavailable reason;
  • extractor name, version, operation, and optional executable digest;
  • coverage status, truncation, omitted count, and named limits;
  • limitations and links to existing Evidence v2 records;
  • a content-derived identifier and its declared stability strategy.

Authority and epistemic state are separate fields. An AST observation does not become a runtime fact, and a passive runtime fact does not prove which static module produced it. In particular:

  • static-relationship-inference can only produce inferred facts;
  • cross-layer-reconciliation can only produce inferred facts and must link both static and runtime Evidence;
  • unknown or unavailable facts require unknown confidence and a limitation;
  • inferred facts require an explicit limitation;
  • exact confidence is reserved for observations;
  • all observations require an actionable location, and artifact-backed facts require a digest;
  • static/native authorities cannot claim runtime locations;
  • truncated coverage must be partial, name the limiting bound, and not claim zero omissions;
  • complete coverage cannot be truncated or omit items.

Unavailable values are represented explicitly rather than as missing fields, so absence of evidence cannot silently become evidence of absence.

Entity identity

Nodes separate stable entity identity from individual observations. Changing a label, extractor, or evidence record changes the observation ID without necessarily changing the node ID.

StrategyStability scopeIntended use
content-digestGlobally exact bytesArtifacts and byte-identical assets
source-map-originalOne source-map commitmentRecovered original source modules
canonical-pathOne artifact versionASAR entries and assets at exact normalized paths
artifact-local-keyOne artifact versionIPC channels, handlers, exports, and other named entities
structural-fingerprintExplicit cross-version inferenceModules whose relationship is hypothesized from bounded structure
runtime-instanceOne capture onlyTargets, frames, runtime scripts, and workers
observation-fingerprintOne observation scopeFacts with no stronger justified identity

Every strategy states its stability scope. Runtime identities commit the capture digest and runtime key; canonical-path and artifact-local identities commit the containing artifact digest. Structural fingerprints require an inferred supporting observation and never claim exact cross-version identity. Content-digest and source-map identities require an observation of the same artifact digest.

Observation IDs use semantic-content-sha256 with observation-exact stability. Edge IDs use the same derivation with relationship-exact stability. Node, observation, edge, and graph IDs are recomputed during parsing, so callers cannot preserve a stale ID after changing semantic content.

Bounds

The v1 boundary admits at most 1,000 roots, 100,000 nodes, 200,000 edges, and 64 observations per node. Each properties object is valid bounded JSON with at most 64 keys, depth 6, 512 structural nodes, and 4,096 characters per string. Arrays of evidence links, limits, and limitations have independent caps.

These are admission limits, not completeness claims. When an extractor reaches a limit, it must record partial coverage and the exact named limit. It must not return a bounded prefix labeled complete.

Domain API

The public domain functions are:

  • createJavaScriptApplicationNode to normalize an entity and derive node and observation IDs;
  • createJavaScriptApplicationEdge to normalize a relationship and derive its ID;
  • createJavaScriptApplicationGraph to sort complete graph content and derive its graph ID;
  • parseJavaScriptApplicationGraph to verify a stored graph and every commitment;
  • serializeJavaScriptApplicationGraph for canonical JSON;
  • computeJavaScriptApplicationGraphSha256 for a byte-stable verified graph digest.

Schema evolution is explicit. A future incompatible shape receives a new schema_version; v1 parsing fails with an unsupported-version diagnostic rather than guessing how to reinterpret it.