JevSQL implementation audit and remediation
September 19, 2026 · View on GitHub
Date: 19 September 2026. Supersedes implementation-audit-2026-09-18.md, which recorded the state before this work.
Scope: the three supplied research documents, re-read in full, compared against the code by reading every module rather than by trusting the previous audit.
Finding
The eight defects the previous audit reported are fixed, each with a regression test that asserts on what leaves the process rather than only on returned rows. Four coverage gaps are closed, and the highest-priority capability named in document A — the typed guardrail in front of an agent-issued statement — is now implemented. The evaluation and operational evidence the documents ask for is still not produced, and is listed explicitly below rather than implied.
Current suite: 378 tests passed, 0 failed, 2 real-server checks skipped locally. Both offline demos run. A live synthetic TypeSafe workflow passed against jev-1.13.0. PostgreSQL and MySQL service tests are configured in CI and are skipped locally when their servers are unavailable.
Source documents
- A: jev by TypeSafe AI: A Typed-Decision Layer for AI-Era Databases — proposal map, pages 4-10.
- B: Jev for SQL Databases: Deep Research, Architecture Map and Three-Month PoC — capability map pages 5-15, control design 16-28, PoC targets 28-31.
- C: JEV Model Database Solutions — proposals, pages 5-12.
Defects fixed
Each was independently reproduced against the code before being changed.
| ID | Defect | Fix | Regression test |
|---|---|---|---|
| F1 | A grouped predicate such as WHERE (tenant_id = 1 AND jev_bool(...)) has no top-level AND, so the whole clause was relaxed for the collect pass and every tenant's text was sent for judgment. Returned rows stayed correct, which hid it. | rewrite.mjs now descends into parenthesised conjunctions. A clause that genuinely cannot be split is reported in stats.widened, and strictCollect: true refuses the query. | "a grouped authorization predicate still narrows which rows are judged" — asserts on the mock server's received payloads |
| F2 | A tenant column named account_id, or TENANT_ID in a case-insensitive catalog, was not recognised, so the compiled SQL carried no tenant predicate while the review evidence still claimed an enforced scope. | query-compiler.mjs requires an explicit classification for every table a tenant-scoped actor touches: a column name, or null for shared. Recognition is identifier-case aware. Evidence now reports the scope actually enforced. | "an unclassified table refuses to compile for a tenant-scoped actor", "a tenant column is recognised however the catalog cases it" |
| F3 | Runbook descriptions and other question criteria travelled to the provider, and into includeEvidence receipts, without passing through redaction. | privacy.mjs adds redactQuestions; decision-service.mjs minimises questions on the same path as state, before hashing or sending. Option labels are preserved because they are part of the answer contract. | "runbook descriptions are redacted before they reach the provider or a receipt" |
| F4 | The proof that a join cannot multiply rows used column names and index uniqueness while ignoring collation, so a NOCASE key joined to a BINARY UNIQUE column produced a summed 20 where 10 was correct. | The uniqueness proof now requires the comparison collation to match the collation that enforces uniqueness; otherwise the join is treated as multiplicative and the aggregate is refused. | "a collation mismatch makes a join multiplicative instead of aggregable", plus the matching-collation case |
| F5 | Denied-column checks read Column opcodes against table_info order. An INTEGER PRIMARY KEY is read with Rowid and produced no column at all; a WITHOUT ROWID table is stored key-first, so a denied column was reported as a different public column. | sql-inspector.mjs derives the real physical layout for WITHOUT ROWID tables, uses index_xinfo for index cursors, resolves Rowid/IdxRowid, and blocks when a read cannot be named while a column policy applies. | "denied columns are detected through rowid aliases and WITHOUT ROWID layouts" |
| F6 | A Choice answer naming an option with probability 0 passed validation and could satisfy a confidence gate. | validation.mjs requires the chosen option to be an argmax within tolerance, matching TypeSafe's definition of Choice. | "a choice that is not its own highest-probability option is rejected" |
| F7 | verify() covered receipts only, so editing a saved human label left it reporting ok: true with an unchanged head hash. | receipts.mjs adds a single _jevsql_chain integrity log covering receipts and labels, backfilled from existing receipts so an archived head hash stays valid. verify() also detects rows written straight to either table. | "the integrity log covers human labels as well as receipts", "a label appended straight to the table, bypassing the log, is detected" |
| F8 | Remote catalog snapshots selected ordinal_position and then discarded it, so a reordered PostgreSQL or MySQL catalog hashed identically and produced no diff. | adapters.mjs preserves position, converting the 1-based catalog value to the 0-based snapshot field and rejecting an invalid one. | "remote-style snapshots keep column positions, so reordering is a change" |
Coverage gaps closed
- Release qualification could pass a policy that blocks everything. Blocking every case gives a perfect false-allow rate. metrics.mjs now measures the safe side too —
safeCases,falseBlocks,falseBlockRate,falseBlockInterval,safeAllowRate— andqualifyReleaseenforcesminSafeCases,maxFalseBlockRateandminSafeAllowRatealongside the existing safety caps. A block-everything policy returnsreview. This matches document B's safety metric family, which lists false-block rate beside false-allow rate. - Measured hash spills were discarded. The
causerubric referred to hash spills that the parser never produced. telemetry.mjs readsHashAgg Batches,Hash Batches,Original Hash Batches,Planned Partitions,Disk UsageandPeak Memory Usage, per worker as well as the leader, exposes them asnode.hash, and raises ahash-spillsymptom on measured disk usage or extra batches. - Column positions in remote snapshots — see F8.
- Documentation understated the control modules. The README now describes both execution paths, their different guarantees, model pinning, the tenant-classification requirement, and the collect-pass widening boundary.
Capabilities added
Document A ranks the typed guardrail proxy first of ten, and document C describes the same pattern as an inline execution firewall. It was the largest missing piece.
DatabaseControl.reviewStatementgates one agent-issued statement against its declared intent.classifyStatementin sql-inspector.mjs decides the operation class, destructiveness and boundedness from masked SQL text, so quoted identifiers cannot fake a keyword and a writable CTE is classified by the write it performs rather than by its leadingWITH. Thestatementreview preset adds intent match, destructiveness, reversibility, shared-data reach, actor expectation, operation class and a blast-radius rubric. Only a read that matches its intent can beeligible; every write needs approval, a destructive, unbounded or permission-changing statement needs an out-of-band human, and a multi-statement batch is refused without a model request. Nothing in this path can execute a statement.- Three further review presets named in document B and absent before:
ormfor already-counted N+1 and model/schema mismatch evidence,costfor grouping measured spend by business purpose and spotting redundant workloads, andsecretsfor the ambiguous prose that regular expressions cannot settle. - The guardrail is exposed as
jevsql control statementand is exercised in the offline control demo.
Second pass: the remaining capability map
A later pass built out the workflows the first pass had listed as missing. Each is deterministic where the documents say it must be, and each is covered by tests.
| Capability | Where | What it establishes |
|---|---|---|
| Migration replay, test packs, rollback verification | migration-runner.mjs | Proves the DDL produces the schema the review was written against, runs the named packs, and checks the rollback restores both shape and rows |
| Application-versus-database type divergence | app-types.mjs | Deterministic comparison of a declared model against the schema, with only the ambiguous residue reviewed |
| Evaluation corpus | corpus.mjs | Document B's per-decision record, with splits derived from the case id and append-only gold labels |
| Shadow scoring and the promotion ladder | shadow.mjs | Scores alongside production without returning anything actionable; reports the highest rung the evidence supports |
| Untrusted-content boundary and adversarial suite | injection.mjs | Structural fencing plus signals that remove eligibility; the bundled suite reaches zero automatic allows with zero false positives on benign controls |
| Lock graphs, backup and replication posture | operations.mjs | Wait-for cycles, RPO/RTO comparison and lag bucketing, all computed before any review |
| Semantic metric layer | semantic-layer.mjs | Metric and dimension selection as bounded decisions, compiled by the existing query compiler |
| Foreign-key-aware synthetic data | seed.mjs | Topological, seeded generation whose rows the database itself accepts |
| Index proposal, measurement, property tests | candidates.mjs | Candidates proposed from access patterns, then built, timed and dropped; rewrites compared on edge-case fixtures |
| Lineage and sensitivity propagation | lineage.mjs | Downstream inheritance that contradicts an optimistic catalogue entry |
| Cascading tiers | escalation.mjs | Measured escalation rate and cost per case, compared against an expensive-only baseline |
Views and triggers now participate in the schema snapshot, hash and diff, so a replaced view or a new insert-blocking trigger is a detected change. Remote snapshots collect indexes, check constraints, collations, generated columns, views and triggers where the server exposes them, and list what it did not in unavailableCatalogs.
What is still not done
These remain open. None is a code-level defect; they are the evidence a deployment would rest on.
- No adjudicated evaluation corpus ships here. Document B proposes roughly 1,000 cases: 400 request/SQL pairs, 250 migrations, 350 incidents. The store, the split discipline, the labelling workflow and the metrics are implemented and tested. The cases do not exist, so every accuracy, calibration, precision, recall, reviewer-time and triage-time target remains unproven.
promotionStatusreports this rather than hiding it: with no corpus, the ladder sits below its first rung. - Real PostgreSQL and MySQL checks need an evidence history. CI provisions PostgreSQL 16 and MySQL 8.4 and tests restricted roles, schema reads, plans, tenant-bound reads, rollback, cancellation, statement timeouts and lock timeouts. The checks are skipped locally when those servers are unavailable. Replay, seeding and index measurement still run on SQLite, so they establish behaviour on the supplied fixtures and the current data volume, not on a production system.
- No shadow pilot has been run. The runner exists; no production traffic has been scored with it.
- Collation is absent from remote snapshots, so the join-multiplication proof falls back to treating both sides as the default collation. That is correct for SQLite and unverified for a MySQL catalog whose columns genuinely differ.
- Lineage discovery from SQL text is lexical. Every edge it proposes is marked
discovered; confirming them is a human step. - Still absent: a live telemetry collector, an ETL orchestrator, a dialect translator (the equivalence gate reviews a translation someone else produced), and partition or shard implementation. Reviews exist for these; the surrounding operational workflow does not.
Claims that should not become guarantees
Document C overstates what constrained output establishes. A schema-valid answer is not a correct one, and F2, F4 and F6 were local counterexamples where a well-typed answer accompanied a wrong result. TypeSafe documents calibration across groups and explicitly separates it from the correctness of any individual answer.
The latency, automation-percentage, cost and zero-error figures in document C are vendor or third-party claims, not JevSQL measurements. Autonomous index creation, failover and unrestricted mutation should not be added to match them. The review-only boundary is deliberate: after this work, the count of database mutations a model answer can trigger on its own is still zero.