Jev: what Typesafe documents, and what fastbrowse assumes

September 18, 2026 ยท View on GitHub

Every Jev assumption in fastbrowse, checked against Typesafe's own documentation (September 2026, Jev 1.13). Anything the documentation does not state is marked ours: a fastbrowse choice, tuned on the evals rather than taken from Typesafe.

The contract

DocumentedWhere fastbrowse relies on it
Direct APIPOST https://api.typesafe.ai/v1/systemone, bearer key, body {model, state, questions}; response {model, answers, usage} (API, OpenAPI)clients/typesafe.py, used when TYPESAFE_API_KEY is set, unless FASTBROWSE_JEV_SOURCE=gateway
GatewayModel typesafe-ai/jev via https://ai-gateway.vercel.sh/v4/ai/evaluation-model, body {state, questions} (Gateway, transport source)clients/vercel.py, used with AI_GATEWAY_API_KEY; FASTBROWSE_JEV_BASE_URL retargets either
Yes/no (Noul)Direct returns {type: "noul", noul: P(yes)}; the gateway returns probability. Optional true/false criteria define the boundary (Noul, v1 migration)clients/validation.py decodes each shape separately
ChoiceThe top choice, every option's probability, and a confidence (Choice)policy.py (operation and target), retrieval.py (field and short-fact reads)
ScoreOrdered levels, returning a probability-weighted index (Score)Modelled in jev.py, not yet called
OptionsAt most 255 per Choice (Choice)MAX_CHOICE_OPTIONS = 255; the 240 cap and group-then-element selection are ours
Tokens32k for state plus the largest question, 64k in total (Models)TokenBudget targets 24k and 48k; the headroom and chars_per_token = 3.0 are ours, because tokens are estimated locally
Rate limits1,200 requests a minute and 250k tokens a second on the direct API, subject to change (Models); no extra gateway limit on paid tiers (Gateway limits)A run makes a few requests a step, far below either
Errors400/401/403/404/422/429/5xx, with 529 for overload; retry with backoff, honouring server retry headers (Exceptions)post_with_retry retries 408, 429, 500, 502, 503, 504 and 529 (not the 4xx request errors) and honours retry-after-ms and retry-after, capped at 10s
Price$0.042 per million input tokens; output is free (Models, gateway catalog)CostComponent.JEV on the ledger
Latency70 to 500ms end to end, as advertised (launch post); no SLA is publishedMeasured through the gateway: 0.28s median and 0.62s worst over 25 policy-sized calls. The 1.5s hedge in clients/validation.py is ours
StreamingGateway evaluation does not stream (AI SDK evaluation)Not needed: answers are a few numbers

Confidence is not correctness

Typesafe says Jev's probabilities are calibrated: across many answers, frequencies match the stated probability (primer). A Choice's confidence summarises how concentrated the distribution is; it is not the chosen option's probability (Confidence). Noul has no separate confidence: its probability is the answer.

Every threshold in config.py (recover_below 0.55, done_accept_from 0.85, claim_problem_above 0.70, the 0.90 read cut in retrieval.py) is ours. None is a Typesafe recommendation. Each is set per question on the evals, and the gates behind them (VERIFY, the claim check, the LLM fallback) mean a miscalibrated threshold tends to cost seconds rather than a wrong answer. Not always: a completion judged at or above done_accept_from skips VERIFY, so a confidently wrong DONE is not caught there.

How the questions are written

Typesafe's guidance (Primitives, Noul, Choice, Jev 1.13 jaggedness), and where fastbrowse follows it:

  • Narrow, explicit judgments, with true/false descriptions that match the instruction. Every Noul question in verification.py and safety.py carries both descriptions.
  • Concrete option boundaries, plus an escape option when coverage is incomplete. The short-fact read offers "none of these", and the policy offers escalate.
  • Question ids are invisible to the model. Everything the model needs is in the instruction text.
  • Batch independent questions that share one state; answers cannot see each other. Each step is one request: the operation, a target for each operation, and whether the page needs a sign-in. The done check asks completion, unmet actions and "does the draft need rewriting" in one call.
  • Remove irrelevant state and keep arithmetic in code. State is the redacted viewport and the notes; counts and comparisons go to the LLM reader.

Not yet used

  • Structured instructions (Structure): jev.py types instructions as strings. Structured criteria are used where they help: the tab choice passes each tab as an object.
  • Score for graded judgments such as relevance.
  • Full Choice distributions for trying a second-best target, rather than only the top pick (hierarchical classification).

No faster Jev tier, cross-request cache or logprob access is documented.