Benchmark Report

July 29, 2026 · View on GitHub

This is the latest checked-in report for patina's deterministic suspect-zone benchmark.

Scope: this benchmark measures whether patina's stylometry layer flags fixture paragraphs as AI-like editing hotspots. It does not prove whether a real document was written by a human or by AI.

Current result

  • Status: passing
  • Generated at: 2026-07-29T08:50:08.132Z
  • Node: v24.18.0
  • Fixture schema: v1
  • Fixtures: 49
  • Languages: 4 (en, ja, ko, zh)
  • Overall accuracy: 100.0% [92.7%–100.0%] (n=49, Wilson score interval, 95%)
  • Source fixtures: tests/fixtures/suspect-zones/**
  • Regression ranges: tests/fixtures/suspect-zones/expected-ranges.json (refresh with npm run benchmark:ranges)
  • Reproduce: npm run benchmark:report
  • Raw JSON: latest.json
  • Detector comparison protocol: detector-comparison.md
  • 2025+ re-baseline plan: docs/research/2025-rebaseline-plan.md

Language breakdown

langfixturesaccuracy95% CIprecisionrecallf1TPFPFNTN
en13100.0%77.2%–100.0%100.0%100.0%17006
ja12100.0%75.8%–100.0%100.0%100.0%16006
ko12100.0%75.8%–100.0%100.0%100.0%17005
zh12100.0%75.8%–100.0%100.0%100.0%16006

Detector breakdown

langdetectorfixturesaccuracy95% CIprecisionrecallf1TPFPFNTN
enburstiness1392.3%66.7%–98.6%100.0%85.7%0.926016
encandor1353.8%29.1%–76.8%100.0%14.3%0.251066
enendingMonotony1346.2%23.2%–70.9%0.0%0.0%00076
enkoDiagnostics1346.2%23.2%–70.9%0.0%0.0%00076
enlexicon1346.2%23.2%–70.9%0.0%0.0%00076
enmattr1346.2%23.2%–70.9%0.0%0.0%00076
enthematicBreak1346.2%23.2%–70.9%0.0%0.0%00076
jaburstiness1291.7%64.6%–98.5%100.0%83.3%0.915016
jacandor1250.0%25.4%–74.6%0.0%0.0%00066
jaendingMonotony1250.0%25.4%–74.6%0.0%0.0%00066
jakoDiagnostics1250.0%25.4%–74.6%0.0%0.0%00066
jalexicon1275.0%46.8%–91.1%100.0%50.0%0.673036
jamattr1250.0%25.4%–74.6%0.0%0.0%00066
jathematicBreak1250.0%25.4%–74.6%0.0%0.0%00066
koburstiness1291.7%64.6%–98.5%100.0%85.7%0.926015
kocandor1241.7%19.3%–68.0%0.0%0.0%00075
koendingMonotony1283.3%55.2%–95.3%100.0%71.4%0.835025
kokoDiagnostics1258.3%32.0%–80.7%100.0%28.6%0.442055
kolexicon1241.7%19.3%–68.0%0.0%0.0%00075
komattr1241.7%19.3%–68.0%0.0%0.0%00075
kothematicBreak1241.7%19.3%–68.0%0.0%0.0%00075
zhburstiness1275.0%46.8%–91.1%100.0%50.0%0.673036
zhcandor1250.0%25.4%–74.6%0.0%0.0%00066
zhendingMonotony1250.0%25.4%–74.6%0.0%0.0%00066
zhkoDiagnostics1250.0%25.4%–74.6%0.0%0.0%00066
zhlexicon1275.0%46.8%–91.1%100.0%50.0%0.673036
zhmattr1250.0%25.4%–74.6%0.0%0.0%00066
zhthematicBreak1250.0%25.4%–74.6%0.0%0.0%00066

Ranking diagnostics

Signal-score ranking shows whether the diagnostic signal_score separates hot fixtures from natural fixtures before any threshold is chosen. It is computed only on the checked-in fixture corpus and is not a broader model-era claim.

scopefixturespositivesnegativesROC-AUCPR-AUCbest thresholdprecisionrecallbest F1accuracy
overall492623113.846100.0%100.0%1100.0%
en13761150100.0%100.0%1100.0%
ja12661123.167100.0%100.0%1100.0%
ko1275113.846100.0%100.0%1100.0%
zh1266116.772100.0%100.0%1100.0%

Low-FPR operating points

TPR at a fixed false-positive budget. Aggregate AUROC/accuracy can hide deployment failure, so these report the strict operating point on the checked-in fixture corpus. n/a marks a slice without enough negatives (or positives) to support the target; max FP of 0 is a strict zero-false-positive point.

scopetarget FPRnegativesmax FPactual FPRTPR
overall1.0%2300.0%100.0%
overall5.0%2310.0%100.0%
en1.0%600.0%100.0%
en5.0%600.0%100.0%
ja1.0%600.0%100.0%
ja5.0%600.0%100.0%
ko1.0%500.0%100.0%
ko5.0%500.0%100.0%
zh1.0%600.0%100.0%
zh5.0%600.0%100.0%

Slice metrics

Report-only confusion metrics grouped by metadata dimension. language, class, and lengthBucket are derived from current fixtures; generator and edited are resolved through the model_family/edit_depth mapper (human controls become generator: human, un-edited rows become edited: none); domain and register default to unspecified until the corpus carries that metadata. Slices below the per-dimension minimum count are reported as insufficient data (counts only). No detector thresholds change.

language (min 5)

valuenaccuracyprecisionrecallf1state
en13100.0%100.0%100.0%1ok
ja12100.0%100.0%100.0%1ok
ko12100.0%100.0%100.0%1ok
zh12100.0%100.0%100.0%1ok

class (min 5)

valuenaccuracyprecisionrecallf1state
ai26100.0%100.0%100.0%1ok
natural23100.0%ok

lengthBucket (min 5)

valuenaccuracyprecisionrecallf1state
medium12100.0%100.0%100.0%1ok
short37100.0%100.0%100.0%1ok

domain (min 5)

valuenaccuracyprecisionrecallf1state
unspecified49100.0%100.0%100.0%1ok

register (min 5)

valuenaccuracyprecisionrecallf1state
unspecified48100.0%100.0%100.0%1ok
workplace-summary1insufficient_data

generator (min 5)

valuenaccuracyprecisionrecallf1state
human23100.0%ok
local-fixture1insufficient_data
unspecified25100.0%100.0%100.0%1ok

edited (min 5)

valuenaccuracyprecisionrecallf1state
none49100.0%100.0%100.0%1ok

Sample sizes

langclassfixtures
enai7
ennatural6
jaai6
janatural6
koai7
konatural5
zhai6
zhnatural6

Misclassifications

All fixtures classified correctly.

Fixture log

fixturelangclassexpectedpredictedoksignalCV bandMATTR bandlexicon/1kKO diagnosticsample lexicon hits
en-ai-01enaihothot80.5120.058 low0.928 high0cold
en-ai-02enaihothot69.8830.09 low0.841 high0cold
en-ai-03enaihothot78.4950.065 low0.828 high0cold
en-ai-04enaihothot76.7170.07 low0.84 high0cold
en-ai-05enaihothot68.9940.093 low0.879 high0cold
en-ai-06-chat-registerenaihothot88.7010.034 low0.814 high0cold
en-ai-07-discourse-candorenaihothot500.358 mid0.872 high0cold
en-nat-01ennaturalcoldcold00.881 high0.898 high0cold
en-nat-02ennaturalcoldcold00.886 high0.884 high0cold
en-nat-03ennaturalcoldcold00.914 high0.882 high0cold
en-nat-04ennaturalcoldcold00.494 mid0.854 high0cold
en-nat-05ennaturalcoldcold00.853 high0.875 high0cold
en-nat-06-single-openerennaturalcoldcold00.552 high0.84 high0cold
ja-ai-01jaaihothot84.9590.045 low0.833 high0cold
ja-ai-02jaaihothot23.1670.23 low0.785 high0cold
ja-ai-03jaaihothot79.0670.063 low0.795 high0cold
ja-ai-04-lexiconjaaihothot1000.56 high0.803 high63.83coldまとめると, 結論として, 重要なのは, デジタル時代において
ja-ai-05-formulaic-summaryjaaihothot1000.155 low0.77 high74.074coldまとめると, 現代社会において, デジタル時代において, 長期的に見ると
ja-ai-06-broad-techjaaihothot1000.243 low0.765 high73.77cold結論として, テクノロジーの進化により, 一方で~他方で, ~と言えるでしょう
ja-nat-01janaturalcoldcold00.487 mid0.719 high0cold
ja-nat-02janaturalcoldcold00.65 high0.796 high0cold
ja-nat-03janaturalcoldcold00.395 mid0.807 high0cold
ja-nat-04-lexicon-coldjanaturalcoldcold00.396 mid0.752 high0cold
ja-nat-05-station-notejanaturalcoldcold00.564 high0.822 high0cold
ja-nat-06-maintenance-logjanaturalcoldcold00.519 high0.88 high0cold
ko-ai-01koaihothot68.9920.093 low0.977 high0cold
ko-ai-02koaihothot75.5450.073 low0.82 high0cold
ko-ai-03koaihothot75.5450.073 low0.79 high0cold
ko-ai-04koaihothot67.3140.098 low0.853 high0cold
ko-ai-05koaihothot67.3140.098 low0.853 high0hot: regular-eojeol-length, low-comma-density, low-suffix-class-diversity
ko-ai-06-chat-registerkoaihothot72.8870.081 low1 high0cold
ko-ai-07-ko-diagnostickoaihothot3.8460.417 mid0.955 high0hot: regular-eojeol-length, low-comma-density, low-suffix-class-diversity
ko-nat-01konaturalcoldcold00.717 high1 high0cold
ko-nat-02konaturalcoldcold00.552 high1 high0cold
ko-nat-03konaturalcoldcold00.68 high1 high0cold
ko-nat-04konaturalcoldcold00.771 high0.975 high0cold
ko-nat-05konaturalcoldcold00.996 high0.998 high0cold
zh-ai-01zhaihothot79.2720.062 low0.902 high0cold
zh-ai-02zhaihothot6.7720.28 low0.734 high0cold
zh-ai-03zhaihothot72.430.083 low0.933 high0cold
zh-ai-04-lexiconzhaihothot1000.748 high0.894 high92.593cold总而言之, 总的来说, 值得注意的是, 在数字时代
zh-ai-05-formulaic-summaryzhaihothot1000.41 mid0.931 high105.263cold综上所述, 值得注意的是, 在数字时代, 从长远来看
zh-ai-06-broad-techzhaihothot1000.443 mid0.912 high90.09cold总而言之, 需要指出的是, 随着科技的发展, 带来了新的机遇
zh-nat-01zhnaturalcoldcold00.506 high0.875 high0cold
zh-nat-02zhnaturalcoldcold00.528 high0.936 high0cold
zh-nat-03zhnaturalcoldcold00.58 high0.907 high0cold
zh-nat-04-lexicon-coldzhnaturalcoldcold00.387 mid0.931 high0cold
zh-nat-05-market-notezhnaturalcoldcold00.598 high0.891 high0cold
zh-nat-06-maintenance-logzhnaturalcoldcold00.549 high0.952 high0cold

How to read this

  • Hot means at least one deterministic signal crossed the benchmark threshold: low burstiness CV, low MATTR, AI-lexicon density, or the conservative Korean diagnostic composite.
  • Cold means the fixture did not cross those thresholds.
  • Signal is the 0–100 diagnostic strength of the strongest deterministic trigger. It supports ranking diagnostics but does not replace the binary hot/cold regression gate.
  • The report is meant for regression tracking and contributor discussion, not for authorship accusation.
  • This deterministic corpus is intentionally small (49 fixtures across en, ja, ko, zh); do not treat 100% fixture accuracy as generalization to new models, genres, or edited AI text.
  • Confidence intervals use Wilson score intervals for the checked-in fixture set; external threshold sweeps and 2025+ model rebaselines are separate research follow-ups tracked in 2025+ Re-baseline Plan.
  • Broader methodology notes live in AI/Human Metrics Research and Quality Checks.