Jev Auto Router

September 20, 2026 · View on GitHub

One Codex session. A fresh model choice for each meaningful call. TypeSafe's Jev chooses from the available GPT model and reasoning-effort pairs; independent verification checks the finished task. Jev Auto Router (Jev Router) is a prototype of that design.

In the architecture, Jev chooses the model and reasoning effort for each call. A local Responses proxy preserves the Codex session and tool loop. Independent verification checks the finished task. Router Compass connects the route, actual usage, and acceptance result to answer one question: Did using less frontier capacity still complete the task correctly and reduce the full delivery cost?

Important

Status: architecture approved; runtime is a validation prototype. This repository has a per-call proxy and tests, but cross-model switching in a real Codex tool loop, the complete verification path, and savings have not been proven end to end. The design below is not a production installation guide.

简体中文 · Architecture · Decision ADR 0017 · Chinese article · License

Prerequisites

  • A TypeSafe account and Jev API key. Get a key through the TypeSafe Quick Start. This proxy reads JEV_API_KEY when it calls Jev. TypeSafe's own examples use TYPESAFE_API_KEY; both variables can hold the same key. Never commit the key.
  • Node.js 22+, a signed-in Codex CLI, and access to the model and reasoning-effort pairs you want to route. A Jev key alone does not make live per-call routing available.
  • This is still a validation prototype. Cross-model switching inside a real Codex tool loop must pass the P0 proof below; there is no production-ready install-and-run path yet.

Before running the local proxy, set the key from your TypeSafe dashboard in the current shell:

export JEV_API_KEY="<your-typesafe-jev-api-key>"

Why route per call?

A coding task mixes hard reasoning about requirements or failures with routine calls to read files, make known edits, and follow up on tool results. Running every call on a frontier model spends scarce capacity on simple steps; running the whole task on a small model risks the difficult ones.

Jev Auto Router makes a choice at each meaningful model call inside the same session. It does not start a new worker for each task. For example, Sol could frame the problem, Luna Max handle explicit follow-up work, Terra resolve an ordinary implementation issue, and Sol analyze a failed test. This illustrates the routing granularity; it is not a measured outcome.

The delivery loop

Jev Auto Router per-call architecture: a Codex session passes through the local proxy, Jev, execution guard, and native Responses before independent task verification

Codex Responses call
   |
   v
local proxy -- OFF? -> host model - infrastructure/verification -> fixed tier - privacy fails -> skip Jev
   |
   v
compact routing state + (model, effort) pairs the host can request now
   |
   v
Jev: one Choice (pinned in production) --timeout/low-confidence/bad answer--> Terra/medium baseline (reason recorded)
   |
   v
native Responses forwarding (events unchanged); record actual model + usage
   |
   +-task end-> independent verification PASS / FAIL -- failure facts return to the same session; Root takeover after two cycles

What Jev does

  • The proxy builds only (model, reasoning effort) pairs that the host can actually request. Jev makes one Choice over those pairs, selecting model and effort together. Local code adds no task-kind table or second semantic router.
  • Jev receives compact state that passes an explicit send check, such as the current step, a bounded tool-error summary, and the current model. The full session still goes through the native Codex model call; raw prompts and tool output are not stored in route logs by default.
  • Production routing pins a validated Jev version; jev-latest is evaluated in shadow. Timeout, low confidence, or an invalid answer falls back explicitly to Terra/medium by default, with a recorded reason. Turning the router off restores the host's original model.

Four capability tiers

TierTypical roleOrdinary calls
Luna MaxExplicit, mechanical follow-up stepsEligible
TerraEveryday work and the fallback baselineEligible
SolHarder reasoning, implementation, and correctionEligible
GPT-6Scarce escalation supported by evidenceIneligible by default

GPT-6 enters the candidate set only with temporary eligibility from a verified reasoning blocker. An explicit user mandate for GPT-6 is enforced as a hard constraint. Luna Max means gpt-5.6-luna/max in the current design; tier names must match model and effort pairs that the host actually supports.

What data goes where

  • Sent to Jev: only the compact state that passes an explicit send check — step type, current model, tool name, exit codes, fixed-length error digests. The full session never leaves the host for routing.
  • Sent upstream: the native session content, exactly as it would flow without the proxy; the proxy never rewrites the response event stream.
  • Stored locally: no raw prompts or tool output in default logs; sensitive content is refused outright rather than truncated, and calls without send eligibility skip Jev for the baseline.

A good route does not prove completion

At the task boundary, independent verification checks the original request against the diff, tests, run results, and any needed semantic judgment. The executing model's claim of success is not evidence. Specific failures return to the same session for correction; after two cycles by default, Root takes over. Verification uses a fixed tier outside the economic routing loop.

Router Compass records Jev's choice, the actual model and effort, usage, cache behavior, latency, and fallback reason for each call, then links those facts to the task's verification result. Unknown usage stays UNKNOWN, never zero. Switching models can lose prompt-cache reuse, so V1 measures the full cost instead of assuming routing saves money.

EvidenceWhat it can establish
Production observationActual model mix, usage, and task verification
Historical replayEstimated prices under explicit assumptions; a counterfactual, not quality proof
Fixed-Terra controlWhether equivalent acceptance and complete Jev/correction accounting show real savings with maintained quality

Current status

The repository contains a prototype local proxy, routing decisions, task verification, Compass data structures, and tests. The P0 gate is a real Codex CLI A→B→A switch within one tool loop across all four tiers, checking authentication, actual model and effort, tool-call IDs, streaming events, cancellation, continuation, and compaction. Until that proof and a controlled comparison pass, this project makes no general savings claim and offers no promise of automatic routing after installation.

skills/jev-auto-router/SKILL.md is the install-and-operate guide; the Skill itself never intercepts model calls — routing happens inside the local proxy.

Try it (prototype)

Requires Node.js 22+. These commands start the local proxy and validate the current code; they do not connect Codex to a production proxy:

npm ci
npm run build
JEV_API_KEY="<your-key>" npm start     # http://127.0.0.1:8787
curl -s localhost:8787/health          # router status, baseline, model catalog
npm test                               # builds, then runs the full suite
npm run typecheck

Once Codex's Responses traffic points at the local proxy, each call returns its route tag through the x-jev-route / x-jev-route-source response headers; GET /decisions shows call and task records. Control signals (task id, step type, forced model) travel as x-jev-* request headers and never enter the forwarded body.

Configuration

Environment variableEffectDefault
JEV_API_KEYTypeSafe Jev API keyrequired when routing is enabled
JEV_ROUTER_OFF1 = kill switch: bypass Jev, restore the host modelunset
JEV_MODEactive | shadow (shadow logs would-be routes, executes the baseline)active
JEV_VERSIONPinned production Jev versionjev-1.13.0
JEV_BASELINEFailure fallback pairgpt-5.6-terra/medium
JEV_CONFIDENCE_FLOORConfidence floor below which calls fall back0.55
JEV_DEADLINE_MSHot-path Jev deadline (tuned from shadow latency)2000
JEV_PORTLocal proxy port8787
JEV_UPSTREAM_BASE_URLUpstream Responses base URLhttps://api.openai.com
JEV_ENDPOINTJev API endpointTypeSafe endpoint

The local proxy entry point is src/index.ts.

License

Apache License 2.0