Jev vs Luna: review classification

September 18, 2026 · View on GitHub

A small, reproducible English benchmark comparing ~typesafe/jev-latest through OpenRouter's Decisions API with openai/gpt-5.6-luna through its chat completions API. Both receive the same review text and the same English classification rubric: topic, sentiment, inferred stars, whether a reply is needed, and whether an actual product defect is reported.

Full-run results — September 18, 2026

JEV vs Luna: field accuracy, median latency, and cost per 1,000 calls scaled from known costs

JEV vs Luna: accuracy by classification field, with API failures counted as incorrect

Download the original PNGs for sharing: overview and accuracy by field. Both are 1800 × 1200 pixels. The cost chart scales the mean reported cost per call to 1,000 calls; it does not represent a separate 1,000-call experiment. Rebuild the charts from the published report with uv run --group charts charts.py (no API calls).

Luna scored slightly higher against the fixture labels; Jev was faster and cheaper per call with a known cost. This is one synthetic diagnostic run, not evidence of a general model ranking.

Executed the benchmark with --repeats 3 --seed 42: 100 reviews × 3 repetitions × 2 models, plus one warmup per model, for 602 completed attempts. The run took approximately 12 minutes (09:31–09:43 UTC). Luna used reasoning effort none and strict structured output; both models used the shared rubric, a 60-second timeout, and no retries.

Quality / reliability metricJevLuna
Valid measured responses299 / 300300 / 300
API errors1 HTTP 5200
Overall field accuracy, failures counted wrong96.13% (1442 / 1500)97.13% (1457 / 1500)
Field accuracy on valid responses only96.45% (1442 / 1495)97.13% (1457 / 1500)
All five fields correct in a response88.33% (265 / 300)90.67% (272 / 300)
Topic accuracy93.33% (280 / 300)93.67% (281 / 300)
Sentiment accuracy93.67% (281 / 300)96.00% (288 / 300)
Inferred rating within accepted range94.67% (284 / 300)96.00% (288 / 300)
Needs-reply accuracy99.33% (298 / 300)100.00% (300 / 300)
Defect accuracy99.67% (299 / 300)100.00% (300 / 300)
Topic macro F10.93070.9194
Sentiment macro F10.93810.9603
Needs-reply macro F10.99491.0000
Defect macro F10.99901.0000

Luna's overall advantage was 1.00 percentage point, or 0.68 points when conditioning on valid responses. Jev's topic macro F1 was higher despite slightly lower topic accuracy: macro F1 gives each class equal weight, while accuracy is influenced by this dataset's unequal class counts. Both models classified defects correctly on every valid response; Jev's missing response accounts for its lower end-to-end defect score.

Latency / cost metricJevLuna
Median latency, valid responses0.647 s1.556 s
Mean latency, valid responses0.675 s1.630 s
p95 latency, valid responses0.910 s2.084 s
Mean latency, all measured attempts0.810 s1.630 s
API-reported cost, including warmup$0.009463 (partial)$0.046290
Attempts with known cost, including warmup300 / 301301 / 301
Mean cost per attempt with a known cost$0.00003154$0.00015379

Luna's median latency was 2.40× Jev's, and its mean known cost per call was 4.88× Jev's. These are observed wall-clock latencies and reported costs for this workload. Jev's failed request took 41.03 seconds and had no reported cost; its total above is therefore incomplete, not a definitive billed total.

Repeatability and label limitations

RepetitionJev field accuracyLuna field accuracy
196.6% (483 / 500)97.2% (486 / 500)
295.2% (476 / 500)97.4% (487 / 500)
396.6% (483 / 500)96.8% (484 / 500)

Jev's HTTP 520 occurred in repetition 2 on rev090; it counted as five incorrect fields. Repetitions reuse the same reviews and are not independent new samples. No statistical significance claim is made.

Several disagreements expose ambiguity in the synthetic labels:

  • Both models classified the wrong-size exchange case (case012) as shipping rather than the expected product topic in all six variant/repetition calls.
  • Both classified the account-login problem (case050) as support rather than the expected other topic in all six calls.
  • Both read the crushed-box complaint that was already resolved (case026) as positive rather than the expected negative sentiment in all six calls.

These are disagreements with the chosen rubric/labels, not necessarily obvious model failures. Labels were not changed after seeing the outputs. The next quality evaluation should clarify these boundaries before running on a separate, independently labeled set of real reviews.

Run evidence

  • Run ID: 20260918T093134.708601Z; manifest status: complete.
  • Requested Jev: ~typesafe/jev-latest; returned model: typesafe/jev-1.13-20260917, provider TypeSafe.
  • Requested/returned Luna: openai/gpt-5.6-luna, provider OpenAI.
  • Portable metrics, confusion matrices, configuration, and source hashes.
  • Local raw responses and full manifest: results/20260918T093134.708601Z/ (ignored by Git).
  • Dataset SHA-256: 79b1e3b16e3582d62a408100e71067c271379a1c05dada0af6e30d7fae595e68.

The report preserves the original run's source filenames and hashes. Since that run, compare_models.py was renamed to benchmark.py and review_task.py to classification.py; the measured results and dataset remain unchanged.

Setup and offline checks

git clone https://github.com/mameli/jev-vs-luna.git
cd jev-vs-luna
uv sync
uv run pytest -q
uv run benchmark.py --dry-run

Install uv first if needed. Offline tests and the dry run do not require an API key. Before running paid API calls, copy .env.example to .env and set OPENROUTER_API_KEY.

The test suite makes no API calls, even if a key is configured. For a paid integration smoke check, use the small comparison below. Never commit .env or credentials.

Run a comparison

# Small paid smoke run: 6 measured calls + 2 warmup calls.
uv run benchmark.py --limit 3 --repeats 1
# Full comparison: 600 measured calls + 2 warmup calls.
uv run benchmark.py --repeats 3 --seed 42

Options include --data, --model, --effort, --timeout, --warmup, --output, and --limit. Zero means no limit; repeats must be positive. Luna uses strict JSON Schema and requires a provider that supports the request parameters. An unsupported model/configuration is reported as an error; it is not silently replaced with a different mode. Reasoning uses OpenRouter's reasoning.effort parameter.

Reference: OpenRouter structured outputs.

Dataset and label policy

data/reviews_100.json contains 50 explicitly labeled synthetic cases, each with two contextual variants (100 distinct texts). The second variant adds neutral order context. Variants share case_id and are not independent samples. Regenerate the file with uv run dataset.py.

Cases cover ordinary feedback, mixed sentiment, sarcasm, negation, historical defects, resolved complaints, and positive/neutral reviews with open questions. Labels are manually specified per case; needs_reply is not derived from sentiment. Historical defects still count after resolution; damage to packaging alone does not count as a product defect. Stars are inferred from sentiment: negative 1–2, neutral 3, positive 4–5. They are not observed customer ratings.

The dataset is intentionally small and synthetic, with unequal class counts. It is useful for debugging and controlled latency measurements, but cannot establish real-world superiority. Review labels independently and add held-out real reviews before drawing broader quality conclusions. Do not tune prompts on the same cases used to claim generalization. No confidence intervals are presented: repeats and contextual variants are correlated observations.

Measurement and saved evidence

  • Each review is sent once to each model per repetition. All repetitions count toward accuracy; failed calls count as incorrect for every field.
  • Review order is seeded and shuffled each repetition. Which model goes first is balanced within each repetition (within one pair for odd dataset sizes). Calls run sequentially with the same timeout and no automatic retries.
  • One warmup per model is the default. Warmups are saved and included in reported cost coverage, but excluded from accuracy and measured latency.
  • Latency is client wall time including request and response validation, not inference-only time. All-attempt and successful-call statistics are separate; p95 uses the nearest-rank definition. An all-failure run remains reportable.
  • Reports include per-field accuracy, micro accuracy, exact match, macro F1 for categorical/boolean fields, and confusion matrices with an error column. Macro F1 excludes labels absent from both expected and predicted values. Rating and sentiment are correlated, so the five-field micro score should not be treated as five independent quality measurements.
  • Cost comes only from numeric API-reported usage.cost in USD. Missing cost is unknown, never zero. Check cost_complete and coverage before comparing totals. Raw usage is preserved for token analysis; no hard-coded prices are used.

Each run creates a unique results/<UTC timestamp>/ directory containing:

FileContents
manifest.jsonConfiguration, models, rubric, selected reviews, source/data hashes, completion status
attempts.jsonlFlushed record of every completed attempt, raw payload, prediction, latency, and error type
summary.jsonAggregated metrics and cost coverage

Interrupted runs preserve completed attempts and have status: incomplete. Their summary covers only recorded attempts and must not be compared as if all scheduled pairs completed. A request interrupted in flight may still be billed without a saved response. Provider routing, caching, and mutable model aliases can change over time; keep raw response metadata when comparing runs. Results may contain review text and provider responses; results/ is ignored by Git. No credentials are written to the manifest.

Files

  • classification.py: shared rubric, output schema, strict validation.
  • jev_client.py: minimal Decisions API client.
  • benchmark.py: paired benchmark and durable reports.
  • dataset.py: explicit English cases and deterministic generation.
  • charts.py: reproducible PNG charts from the published benchmark report.
  • assets/: chart images for the README and social posts.
  • tests/: offline regression coverage for the benchmark.