jev-router for Claude Code
September 20, 2026 · View on GitHub
Two pieces, sharing ~/.omp/agent/jev-router.json with the OMP extension:
proxy/— a gateway model calledjev-router. Per-turn model + effort routing with Jev, done at the request layer because a hook cannot do it.hooks/— aPreModelSwitchcache guard. Reports (or refuses) what a manual/modelswitch re-sends uncached.
Why a model and not a hook (checked against the docs)
- Hooks cannot switch the model.
PreModelSwitchcanallow,ask, ordenya switch that someone else requested;set_modelcomes from "an Agent SDK host or Remote Control", not from a plugin. - A gateway can. Claude Code sends every request to
ANTHROPIC_BASE_URL; behind a gateway "your provider or gateway defines the model names, so Claude Code passes any string through without checking it". modelOverridesgives the alias a real model's identity. "To give a gateway alias the capabilities of the model behind it, map that model's Anthropic ID to your alias with amodelOverridesentry." So the settings carry{"modelOverrides": {"claude-opus-5": "jev-router"}}and you selectclaude-opus-5: Claude Code uses Opus's window, tool search and picker label, and sendsjev-routeron the wire. (ANTHROPIC_CUSTOM_MODEL_OPTIONalso works but "_SUPPORTED_CAPABILITIES… have no effect behind anANTHROPIC_BASE_URLgateway", and the alias is then an unknown model — see the 134k-token finding below.)ENABLE_TOOL_SEARCH=1keeps MCP tool deferral on behind the gateway.- Your login survives. "Setting only
ANTHROPIC_BASE_URL, without a gateway credential, doesn't replace the subscription" — the proxy forwardsAuthorizationandanthropic-betaverbatim, which the OAuth path requires.
The gateway model
bun add -g github:devjtv/jev-router # or `bun link` in a clone → `jev-router` on PATH
jev-router setup # guided onboarding; --yes for defaults
setup walks: detect (bun, claude, key, daemon) → OpenRouter key, verified with
one real Jev call → a model per tier (fast / standard / deep / planner) → what
Claude Code should believe it runs (Opus 5 / Sonnet 4.6 / Fable 5.1 — always a bare
id; pick the 1M row in /model afterwards, which the override still matches) → shadow mode, subagent
policy, cache guard → writes jev-router.json (merge) → starts the daemon →
optional login service → wires settings.json (merge, backup) → status bar
choice: if you already have a statusLine, keep yours with the jev line
chained above it, replace it, or leave it → two live dry runs.
The same, as individual commands:
jev-router start # background proxy (pidfile + log in ~/.omp/agent)
jev-router env --write [--statusline=if-absent|replace|chain|skip]
# merge env + modelOverrides (+ statusline) into ~/.claude/settings.json
jev-router service install # start at login (systemd --user / launchd / Task Scheduler → Startup folder fallback)
jev-router update # pull latest (git) or re-add (bun add -g), bun install, restart the daemon
jev-router claude [args] # one session on the gateway model; reuses the daemon or runs a private proxy
jev-router status | stop | restart | reload | logs -f | route "<prompt>" | serve
bun claude-code/launch.ts [args] is the same as jev-router claude. CLAUDE_BIN
overrides which claude runs (a real executable is preferred over a .cmd
shim on Windows; an npm-left claude.cmd on this machine is 160 zero bytes and
exits silently).
Any OpenRouter model, per tier
jevr models # tiers, and whether an OpenRouter key is set
jevr models gemini # search ~450 models ("qwen coder" works too)
jevr models --tier fast --pick # autocomplete → model → provider variant
jevr models --tier fast --set openrouter/qwen/qwen3-coder-flash --variant nitro
jevr is the short form of jev-router (bare jevr launches Claude Code).
Specs are openrouter/<their id>, plus OpenRouter's provider variants as
suffixes: :nitro (throughput, priority tier) or :floor (cheapest, flex
tier). OpenRouter serves the Anthropic Messages format, so nothing is
translated; the proxy swaps in the OpenRouter key and never sends the claude.ai
token there. Anthropic-direct tiers keep using your subscription, so a single
config can mix both.
provider.sort/only/ignore/allow_fallbacks/zdr live under
claudeCode.openRouter; a :nitro/:floor suffix already implies the sort,
so it is not stacked.
Measured, not assumed: Gemini 2.5 Flash :nitro drove a full Claude Code
turn — the model called Read, Claude Code executed it, the answer was right.
A tool_format_suspect line appears (log + daemon log, once per model) if an
upstream returns 200 with a tool call written as text instead of tool_use,
which Claude Code would show rather than run: measured once with Qwen3 Coder
Flash through the provider that served it, while a minimal probe of six models
(Claude Haiku, Gemini Flash, GPT-4.1-mini, DeepSeek, Qwen3 Coder ±flash)
returned proper tool_use — so the variance is per provider as much as per
model. Deferred tools (defer_loading) are an Anthropic-only capability and
are stripped up front for non-Anthropic models, which OpenRouter rejects them
for.
Daemon details
- The server writes the pidfile (
~/.omp/agent/jev-router.pid) with its bound port, so"port": 0works.statusreports running only when the pid is alive andGET /jev-router/statusanswers with that pid; anything else is "not running" andstop/startclear the stale file. startspawnsjev-router serve --quietdetached (windowsHide), with stdout/stderr in~/.omp/agent/jev-router-proxy.log, and polls readiness for up to 10 s. On Windowsstopusestaskkill /Tsince Bun delivers no SIGTERM there.reload(orSIGHUP) re-readsjev-router.jsonin place; sessions and the learned per-model field quirks survive. Sessions idle for six hours are evicted.service installwrites a systemd user unit, a launchd agent, or a Task Scheduler logon task that runsjev-router start;service showprints it first. On Windows the task's console flashes once at logon.env --writemerges only what it owns intosettings.json(CLAUDE_CONFIG_DIRrespected):env.ANTHROPIC_BASE_URL,env.ENABLE_TOOL_SEARCH,modelOverrides.<behavesAs>; setsmodelandstatusLineonly when absent; writes a.bak; refuses a file that does not parse.jev-router claudepasses the samemodelOverridesvia--settingsfor one session instead.
Seeing where a turn went
jev-router statusline is a Claude Code statusLine command: it reads the
session id Claude Code hands every status refresh, asks the daemon, and prints
jev ▸ fast → claude-haiku-4-5 (low) │ 1 subagent │ ctx 20%
jev ▸ waiting for first turn │ ctx 5%
jev ▸ proxy not running
env --write installs it if you have no statusLine; otherwise add
jev-router statusline to your own script's output. The session id in the
statusline payload is the same value as the x-claude-code-session-id header
the proxy keys on (checked live), so a subagent count and the pinned tier are
exact, not inferred. jev-router logs -f and jev-router status show the same
from the terminal.
Or let your agent do it — paste into Claude Code (or any agent with a shell):
Set up jev-router for Claude Code (https://github.com/devjtv/jev-router): run `bun add -g github:devjtv/jev-router` (if that fails: git clone to ~/.jev-router, `bun install`, `bun link`), then run `jev-router setup` interactively in my terminal and let me answer its prompts — it asks for an OpenRouter key if none is found, picks tier models, starts the daemon, wires ~/.claude/settings.json with a backup, and ends with a live dry run. If you cannot run it interactively, run `jev-router setup --yes` and show me its output, then `jev-router status`. Do not edit ~/.claude/settings.json or ~/.omp/agent/jev-router.json by hand.
What happens to a request
| request | proxy |
|---|---|
any path, model ≠ jev-router | forwarded byte-for-byte (background Haiku traffic, subagents with their own model:) |
POST /v1/messages, no tools or max_tokens ≤ backgroundMaxTokens | housekeeping (titles, summaries): first tier's model, no gate call, session pin untouched |
new user turn (last non-system message is user text) | prompt text → Jev → tier → models[tier], output_config.effort from the candidate; pinned to the session |
same turn again (retry, or tool_result continuation) | pinned model and effort, no gate call |
POST /v1/messages/count_tokens | model rewritten to the pinned model |
HEAD /api/hello, GET /jev-router/status | answered locally |
Sessions are keyed by x-claude-code-session-id plus x-claude-code-agent-id
for subagents. The turn fingerprint hashes the user's own text with Claude
Code's <system-reminder> blocks removed, so a re-send with different injected
context is still the same turn; a trailing role: "system" message (per-message
output_config) is skipped when looking for the user's message — Claude Code
sends one on every main-loop request today.
Guards, in this host
- Cache guard: context size is the upstream's own
usage(input+cache_read+cache_creation) from the previous response, read off the SSEmessage_startevent; the first turn uses a chars/4 estimate. AbovecacheGuardTokensa cross-model route keeps the model and changes only effort. - Thinking signatures: when a turn changes the model, prior assistant
thinkingblocks are dropped (stripThinkingOnSwitch) — a signature from one model is not valid on another. Inside a turn the model never changes, so nothing is stripped. - Field compatibility: a 400 naming
output_config/effort,thinking, orcontext_management/clear_thinkingstrips that field and retries once; the model is remembered so later requests are pre-stripped. Droppingthinkingalso dropsclear_thinking_*context edits, which the API rejects without it. - Images:
onImages: "skip"sends the turn tofallbackModelwithout a gate call;"model"usesvisionModel;"route"gates it like text. - Shadow:
shadow: truelogs the would-be route and keeps the current model. - Fail open: gate error →
fallbackTier; proxy error → a 502 withx-should-retry: true; a continuation the proxy has no memory of (restart mid-turn) goes tofallbackModel.
Live run
Against the real Jev gate and the real API through a claude.ai OAuth login (Claude Code 2.1.268, Windows):
background claude-haiku-4-5 toolless=true "<session>…Write the title…" (no gate call)
route fast → claude-haiku-4-5 conf=0.99 579ms "what is 59*7? number only"
compat claude-haiku-4-5 rejected effort → retried without; rejected adaptive thinking → retried without
route deep → claude-opus-5 conf=0.46 248ms "Read package.json … refactor how config is loaded …"
continuation ×7 claude-opus-5 (pinned) count_tokens → claude-opus-5
Both prompts answered correctly.
Second live pass — the 190k-tokens-at-start bug. With the alias installed
as an unknown model, an interactive session showed 134k tokens of baseline
prompt where a direct Opus session showed 53k. Cause: behind the gateway Claude
Code inlined every MCP tool schema (151 tools, 0 deferred, no tool-search beta)
and assumed a 200k window. Fix, measured in the same setup: modelOverrides
(so the session is Opus to Claude Code) + ENABLE_TOOL_SEARCH=1 → 22 tools,
4 deferred, 39.8k tokens (20%), no catalog warning. Two more things fell out
of that pass and are now handled: a 200k plan rejects the context-1m beta
when a turn is routed to Haiku (the proxy strips the header and retries), and
the cache guard used to fire on a session's first turn from a size estimate
alone, pinning Opus — there is no cache to protect on a first turn, so it no
longer does. Remaining rough edge: /cost bills the alias at Opus rates
regardless of the tier that ran; use jev-router logs or /jev-router stats.
Third live pass — the [1m] silent-unrouted bug. [1m] is a request-time
modifier Claude Code appends after resolving a model id (it also brings the
context-1m-* beta, which is why the compat layer strips that beta per model).
It is not part of the id, so a modelOverrides key of claude-opus-5[1m]
matches nothing: Claude Code then puts a name the proxy does not answer to on the
wire, body.model !== cc.model sends the turn to passthrough, and nothing is
routed — with no error from either side. Measured: select claude-opus-5[1m]
with the override keyed claude-opus-5 → routed (fast → claude-haiku-4-5,
answer correct); key it claude-opus-5[1m] → zero routing and, before this fix,
zero signal. Three changes: baseModelId() normalizes behavesAs on read (old
configs self-heal), claudeSettings() emits the base key and base model
whatever the config says, and the proxy compares the incoming model on its base
id so a jev-router[1m] wire name still routes instead of falling through. The
silence is now a warning: after 5 requests with none for the alias, the daemon
logs alias_never_seen and jev-router status prints what to select. status
also says when routing is switched off ("enabled": false), which is invisible
for the same reason.
Config
claudeCode in ~/.omp/agent/jev-router.json — see the root README for every key.
The tier → model map defaults to Haiku 4.5 / Sonnet 4.6 / Opus 5; a tier with no
entry derives its id from an anthropic/<id> candidate in the OMP tier config.
JEV_ROUTER_DEBUG=1 additionally logs every inbound request, passthrough, and
classification, which is how the trailing-system-message shape was found.
The PreModelSwitch cache guard
When you switch models by hand and the cache is warm, this hook tells you what that switch re-sends uncached — and can ask for confirmation or refuse it.
PreModelSwitchruns command, http, and mcp_tool hooks only — notpromptoragenthooks — so the decision here is rules-only. That is fine: the hook payload carriescontext_tokensandprompt_cache_warm, which are the only facts the decision needs, and it means zero added latency.- Claude Code's own docs make the same cost point this plugin exists for: "Each model has its own prompt cache, so the first request after a switch re-reads the whole conversation uncached."
- A PreModelSwitch hook that does not answer before its timeout blocks the
switch, so this hook always answers, and every error path answers
allow.
Install
# one session, no install
claude --plugin-dir C:/Users/jakey/Repos/jev-router/claude-code
# or as a marketplace
claude plugin marketplace add C:/Users/jakey/Repos/jev-router/claude-code
claude plugin install jev-router@jev-router
The plugin root ships its own catalog at .claude-plugin/marketplace.json, so
pointing a marketplace at the claude-code/ directory installs the plugin in it.
Behaviour
Driven by cacheGuardTokens and cacheGuardMode in
~/.omp/agent/jev-router.json (or JEV_ROUTER_CONFIG):
cacheGuardMode | above cacheGuardTokens, cache warm |
|---|---|
off | allow, silent |
effort-only (default) | allow, and report the uncached re-send to the user |
same-family | allow, and report (this host cannot see model families) |
keep | deny with the token count |
Below the threshold, or with a cold cache, or with no token figure in the
payload, the hook stays silent and allows. JEV_ROUTER_CC=allow|ask|deny forces
the action above the threshold while keeping the explanation.
Default threshold: 60,000 tokens. Default mode reports rather than interferes — a switch the user asked for is the user's call, and the useful contribution is the price, not a veto.
Output
{
"systemMessage": "jev-router: 212,431 tokens in context will be re-sent to claude-opus-5 uncached.",
"hookSpecificOutput": {
"hookEventName": "PreModelSwitch",
"permissionDecision": "allow",
"permissionDecisionReason": "212,431 tokens in context will be re-sent to claude-opus-5 uncached."
}
}
Verifying
bun test test/claude-code-proxy.test.ts # 33 tests: request shaping + the proxy against a stub upstream
bun test test/claude-code-hook.test.ts # 15 tests: the hook as a spawned command + plugin structure
bun test test/cli.test.ts # 8 tests: pidfile, settings merge, service definitions, a real start→reload→stop cycle
The proxy tests drive createProxy the way Claude Code drives a gateway —
POST /v1/messages with Claude Code's headers — against a stub upstream that
records what it received and answers JSON or SSE: passthrough of other models,
a routed turn, pinning across retries and tool calls, re-routing on the next
prompt with thinking blocks stripped, the cache guard demoting to effort-only,
compat retries for effort/thinking/context_management, an upstream 529
relayed with its retry headers, gate failure, shadow mode, subagent policies,
image turns, count_tokens, housekeeping requests, and the local endpoints.
The live run above is the end-to-end check.
Eleven tests run the hook the way Claude Code does — as a command reading JSON
on stdin — covering cold cache, small context, missing token count, the
threshold boundary, each mode, the env override, malformed input, an
unresolvable config path, and an empty stdin. Four more assert the plugin
structure: manifest and catalog parse and agree, hooks.json points at a file
that exists after ${CLAUDE_PLUGIN_ROOT} substitution, and the script is
runnable by shebang.
The hook is not verified end to end inside a live session; it was exercised at the process boundary, and its event names, payload fields, and output shape come from the published hooks reference.
Files
proxy/routing.ts pure request shaping (turn detection, prompt text, tier -> model, field stripping)
proxy/server.ts the gateway model: Bun.serve, per-turn routing, streaming relay, usage tracking
proxy/daemon.ts pidfile lifecycle, detached start/stop, service adapters, settings.json merge, claude launch
../bin/jev-router.ts the CLI
launch.ts thin wrapper ≡ `jev-router claude`; --env, --tail
.claude-plugin/plugin.json hook plugin manifest
.claude-plugin/marketplace.json catalog, so the dir can be added as a marketplace
hooks/hooks.json registers PreModelSwitch with a 10s timeout
hooks/pre-model-switch.ts the guard (reuses the OMP extension's cacheGuard)