sepia

September 22, 2026 · View on GitHub

English | 繁體中文 | 简体中文

behavioral eval version consistency release license: MIT

De-AI writing at the layer that actually gives AI away. Fiction gets its narrative architecture repaired before anyone touches word choice; professional documents (release notes, PR replies, postmortems, tickets, technical articles) each get rules matched to their venue.

A portable Agent Skill: any agent that speaks the standard can load it, and the Skills CLI, which supports 77+ agents, installs it with one command. Claude Code, Codex, Grok Build, Antigravity, and QwenPaw additionally get native plugin packaging. One canonical SKILL.md, no per-platform forks. Four operations: write, review (diagnose only), refactor (minimal edits), recreate (full rewrite).

Table of Contents


Why another humanizer

Every popular humanizer edits word choice and syntax. StoryScope (Russell et al., 2026: 61,608 stories, human + 5 frontier LLMs) showed that a classifier using narrative-structure features alone detects AI fiction at 93.2% macro-F1. In the same study's LAMP-edited condition, where human editors had rewritten the surface style, detection dropped only from 95.5% to 93.9%. The tells that survive are architectural: themes explained by the narrator, single-track causally-tidy plots, emotions rendered only as bodily sensation, no real-world references, no reader, linear time, endings resolved by protagonist growth and acceptance.

sepia turns those measured gaps, together with the related studies digested in research/, into a three-pass writing and revision protocol for fiction:

PassLayerExamples
1Narrative architecture (fiction)stop explaining the theme, loosen the causal chain, back-load revelations, mix emotion modes, sparse character networks, name real things
2Discourse flowde-template the paragraph-question sequence, fix the mid-story sag, vary rhythm and positions
3Surface stylethe classic layer: clichés, syntax templates, vocabulary, register

A 30-feature diagnosis rubric and per-model fingerprints across two layers apply when the writing or executing model is known:

Model familyNarrative layer (StoryScope)Sentence-level prose layer (Vendor prompting guides)
ClaudeMeasuredClaude Fable 5.1 and Mythos 5.1, Fable 5 and Mythos 5, Opus 5, Opus 4.8
GPTMeasuredGPT-5.6, GPT-6 Astra
GeminiMeasuredGemini 3 and 3.1
DeepSeekMeasuredConsulted (no guidance published)
KimiMeasuredConsulted (no guidance published)

Notice: Vendors that publish no prompt guidance are recorded as consulted, not guessed.

Professional prose fails differently, and the structure-level finding holds there too: a 2026 replication of StoryScope on 2,250 company blog posts against 11,250 AI mirrors separated them at 98.0 macro-F1 from structural features alone, with the AI shape described as tidy and self-announcing (SLOPSHAPE-2026 in the ledger, arXiv:2609.15369; a preprint whose features are LLM-scored; it tested detection of original and model-self-reworded posts, never human editing). The studies digested in research/ point at filler that carries no information, hedging where a judgment was needed, chatbot leftovers, register that ignores the venue, and formatting that looks stamped out. Each document type gets a thin rule file on top of one shared checklist:

DomainThe gist
Release notes / announcementsuser impact first, artifacts per claim, no marketing inflation
PR / issue repliesanswer first, cite file:line, no reflex praise, length ∝ stakes
Postmortemsblameless toward people, merciless toward mechanisms; timestamps, dead ends, owned action items
Tickets / work orderstitle = outcome, testable acceptance criteria, link don't repeat
Technical articlesopen at the problem, one real dead end, one committed opinion, numbers with conditions
Long-form journalism (features, investigations, data stories)lead and body in two registers, quotations keep their spoken texture, every number carries a comparison, no summary ending

Governing principle: Calibrate to the human distribution, don't invert the AI one. Humans sit at moderate values; a story with every rule applied is a new fingerprint. The skill selects 3–5 moves per story and leaves slack.

Operation entries

The complete plugin package gives Claude Code, Codex, Grok Build, and Antigravity a general router plus five direct entries. QwenPaw gets the /sepia router only, so the table below does not apply there:

OperationClaude CodeCodexGrok BuildAntigravityMeaning
write/sepia-write$sepia-write/sepia-write/sepia-writeCreate new prose
review/sepia-review$sepia-review/sepia-review/sepia-reviewDiagnose without editing
refactor/sepia-refactor$sepia-refactor/sepia-refactor/sepia-refactorMake minimal in-place edits
recreate/sepia-recreate$sepia-recreate/sepia-recreate/sepia-recreateRewrite from the source facts and intent
hemingway/sepia-hemingway$sepia-hemingway/sepia-hemingway/sepia-hemingwayWrite or refactor fiction with the built-in Hemingway voice applied

The general /sepia (Claude Code, Grok Build, Antigravity, and QwenPaw) or $sepia (Codex) router remains available; on QwenPaw the package installs the six skills into each workspace and registers no per-operation slash commands. What was verified on each platform is stated under Install.

Notice: Standalone wrapper installation is unsupported. The operation wrappers depend on their sibling canonical skill; install the complete plugin package.

Experimental: composing with voice skills

Since v0.4.0, sepia defines an interface for stacking a voice or style skill on top of it — a minimalism method, a brand voice, a persona guide. It is opt-in: tell sepia the voice skill is in play, and it loads references/voice-skills.md over the normal route. No external voice is loaded unless you say so.

The interface contract operates under fixed precedence rules across operations and routes:

Rule dimensionContract specification
Architecturesepia's architecture decisions come first.
Move selectionVoice moves are applied selectively (3–5 signature moves per piece, fewer when a sparse shape or the facts offer fewer; formula endings deliberately broken sometimes).
Review diagnosticsReview reports the voice's known costs instead of fixing them away.
Uniformity enforcementUniformity findings keep full strength: a voice does not excuse a metronome.
Professional registerOn professional routes, the venue still sets the register.
Conflict resolutionDirect conflicts come back to you.

Notice: The voice-skills interface is grounded in one blind review experiment on a strict-minimalism specimen — a worked example, not measured evidence.

Two built-in profiles ship under references/voices/, alongside a persona profile kind with declared override rights:

ProfileTarget routesSource & characteristicsOpt-in & invocation
Hemingway (references/voices/hemingway.md)Fiction and professional proseIceberg omission for fiction, the Kansas City Star rules for professional prose; each move traced to its sourceDirect entry /sepia-hemingway (or $sepia-hemingway). On fiction, asking for strong de-AI on a story counts as opting in (sepia announces the profile and how to decline). Without an opt-in, a fiction review only reports when the text fits the profile and loads nothing.
Taiwan long-form journalism (references/voices/tw-journalism.md)Professional routes onlyNine narrative shapes with their moves, drawn from private human-side journalism measurement and its close readingOpt-in phrase: "apply the Taiwan journalism voice". Follows the same rules as Hemingway.
Persona template (references/voices/PERSONA-TEMPLATE.md)Author-definedOne writer's style with declared override rights; checked by scripts/check_persona.pyAuthor-created profile. Template at references/voices/PERSONA-TEMPLATE.md.
Nyaneko (references/voices/personas/nyaneko.md)Professional routes onlyBuilt-in companion voice written from the maintainer's voice specification with her own exemplarsOpt-in phrase: "apply persona Nyaneko" or 「套用 persona Nyaneko」.

Writing with a persona: a short how-to

sepia's rules remove tells. Warmth has another source: the writer knowing who is speaking, and to whom. One case from this project's own release, with nothing measured: the v0.12.0 Threads post was drafted twice by the same model from the same release notes. The first prompt handed it a skeleton (version line, bullets, fixes, engineering note, update line) and got a form filled in. The second prompt handed it an identity, "I maintain this project and just shipped v0.12.0; talk to my followers in my own voice; how you organise it is yours", plus the notes as the only fact source and the platform's limits, and got the post the maintainer kept. A sepia review of that post then found its tells clustered at the opening and the close (an opening about how the change came to be, a maxim ending, a few absolutes), and a refactor took them out without touching the middle.

The order that worked:

  1. Identity first.
  2. Facts separately, as the only source.
  3. Constraints stated as platform limits, with the structure left to the model.
  4. sepia as the check.
Do not use tools; everything you need is below.
I am <who>, writing for <whom>, about <what just happened>. Use my own voice; how to organise it is yours.
Platform limits: <length, markup, mentions>.
Facts come only from the notes below; numbers and identifiers must match them.
===== notes =====
<the source document>

A persona profile is that identity written down once so every piece starts from it: what they are to the reader, what their first sentence does, how warmth attaches to a fact, what they never do, and pieces in their own voice as exemplars. The template is references/voices/PERSONA-TEMPLATE.md, the built-in references/voices/personas/nyaneko.md is a worked example, and the interface validator checks the profile structure:

python3 scripts/check_persona.py <file>

Write your own; the exemplars carry more than the description does.

Sentence rhythm and Chinese calibration

The style pass checks the spread of sentence lengths, the one syntactic measure on which every study that measured it agrees: human text varies more within a passage, in English and in Chinese.

Syntactic measureStatusEmpirical basis
Sentence length spreadChecked signalWithin-passage variance is consistently higher in human text across English and Chinese studies. Evidence and numbers are in research/rhythm-syntax.md.
Mean sentence lengthDiscardedMeasured directions contradict each other across corpora.
Punctuation countsDiscardedMeasured directions contradict each other across corpora.
Paragraph lengthDiscardedMeasured directions contradict each other across corpora.

Chinese text loads references/languages/zh.md, calibrating the style pass across four distinct evidence layers:

Layer / SourceNature of evidenceCorpus & calibration details
HC3 (2023)Measured corpusHuman-vs-machine Chinese corpus baseline.
Traditional Chinese journalismHuman-side empirical measurementAbout two thousand articles from one unnamed Taiwanese publication spanning about ten years (corpus not distributed; digest in research/zh-news-corpus.md).
Contrast groupMachine-side empirical contrast119 synthetic pieces across three models using one shared base prompt with a one-line variant for one model.
Taiwan conventionsNormative standardShort normative section of Taiwan writing conventions drawn from public standards and first-tier consensus.

Notice: The limits of all three empirical Chinese sources are stated in references/languages/zh.md.

Install

Notice: Every command below is written for user scope — install once, use it in every project.

Any agent (Skills CLI, 77+ agents)

npx skills add Nanako0129/sepia -g     # -g = user scope; the default is project
npx skills update sepia -g             # update
npx skills remove sepia -g             # uninstall

Installs on every agent the Skills CLI supports — Cursor, Cline, Windsurf, Copilot, OpenCode, goose, and more. Pick your agents when prompted. Runtime behavior outside the five platforms below has not been exercised by us; the skill is plain markdown under the Agent Skills standard, so file an issue if your agent trips on it.

Notice: "Verified" means the install completes and the sepia entries appear. The five native plugin installers were each exercised with a live install (QwenPaw's by its contributor). Whether the entries then behave as documented has not been checked platform by platform.

Claude Code

# install
claude plugin marketplace add Nanako0129/sepia
claude plugin install sepia@sepia --scope user

# update
claude plugin marketplace update sepia
claude plugin update sepia

Tip: The in-session /plugin install dialog asks you to pick a scope — choose User there.

Codex

# install
codex plugin marketplace add Nanako0129/sepia
codex plugin add sepia@sepia

# update — refresh the marketplace snapshot, then re-add to pick up the new version
codex plugin marketplace upgrade sepia
codex plugin add sepia@sepia

Grok Build

# install
grok plugin install Nanako0129/sepia --trust

# update
grok plugin update

Grok also auto-discovers a Claude Code install of sepia if you have one; either route works.

Antigravity

# install directly from GitHub
agy plugin install https://github.com/Nanako0129/sepia

QwenPaw

# install: qwenpaw takes a local directory (or a zip URL), so clone first;
# the package's skills symlink resolves inside the clone
git clone https://github.com/Nanako0129/sepia
qwenpaw plugin install ./sepia/.qwenpaw-plugin

# uninstall
qwenpaw plugin uninstall sepia

Notice: Contributor-verified on QwenPaw 2.2.1 (#250, not reproduced by the maintainer): the install completes and /sepia is routed, with the packaged skills symlink followed into a real tree by shutil.copytree.

Project scope (alternative)

When one repo should pin its own copy, commit skills/sepia/ into that repo as .agents/skills/sepia (Codex + Antigravity) or .claude/skills/sepia (Claude Code).

Uninstall

Each tool uses its native command:

# Claude Code
claude plugin uninstall sepia@sepia --scope user

# Codex
codex plugin remove sepia@sepia

# Grok Build
grok plugin uninstall sepia

# Antigravity
agy plugin uninstall sepia

# QwenPaw
qwenpaw plugin uninstall sepia

Layout

sepia/
├── plugin.json              # Antigravity packaging
├── skills/
│   ├── sepia/                # canonical skill (Agent Skills standard)
│   │   ├── SKILL.md          # routing, operations, calibration rules, guardrails
│   │   └── references/       # passes, rubric, fingerprints, domain rules, languages/zh.md, voice-skills (experimental)
│   ├── sepia-write/SKILL.md  # thin fixed-operation wrappers
│   ├── sepia-review/SKILL.md
│   ├── sepia-refactor/SKILL.md
│   ├── sepia-recreate/SKILL.md
│   └── sepia-hemingway/SKILL.md  # fiction write/refactor with the built-in voice
├── .claude-plugin/          # Claude Code packaging (plugin.json, marketplace.json)
├── .codex-plugin/           # Codex packaging
├── .qwenpaw-plugin/         # QwenPaw packaging (plugin.json, plugin.py, skills symlink)
├── .agents/                 # Codex/Antigravity workspace-mode discovery + Antigravity workflow
└── research/                # digested evidence base with sources

Star History

Star History Chart

Star history growth over time for Nanako0129/sepia.

Sources

Full digests with links are in research/. Primary studies include:

SourceVenue / Identifier
StoryScopearXiv:2604.03136
LAMPCHI 2025
Measuring AI SloparXiv:2509.19163
Reinhart et al.PNAS 2025
Russell et al.ACL 2025
NarraBencharXiv:2510.09869
Echoes in AIPNAS 2025
QUDsimCOLM 2025
Beguš2024
Beyond CheckmateEMNLP 2025
Nonaka & Perry2025
Chakrabarty et al.2026
Shan, Lee & Hao2026
Rohrbacher et al.2026
Sourati et al.2026

Support

You can use sepia for free without an account. The research behind every rule is open. Ongoing costs are just maintainer time and two kinds of model quota: delegating literature surveys to research agents that read primary papers, and running live models to test rule changes on A/B stories and cross-platform end-to-end reviews before shipping. You can support the project on Patreon.

Support sepia on Patreon

License

MIT