falsegreen-skill on OpenAI / Codex
July 30, 2026 · View on GitHub
How to use the falsegreen-skill J1-J6 protocol with ChatGPT, the OpenAI API, structured output, Codex CLI, and batch pipelines.
Installation (Codex CLI)
Codex has no single-command install for this repo (see the marketplace note below). There are two working paths; pick by whether you want the full protocol or a persistent hook in your own project.
Full protocol (clone). Cloning gives Codex every protocol file with its
relative references intact, so on-demand escalation to reference.md and
SKILL.md resolves. AGENTS.md at the repo root auto-loads when you start
Codex inside the clone:
git clone https://github.com/vinicq/falsegreen-skill
cd falsegreen-skill
codex
The protocol is scoped to this directory: when you cd to your own project,
Codex no longer has it loaded. The shared skill lives at
skills/falsegreen-skill/SKILL.md and the plugin manifest at
.codex-plugin/plugin.json.
Persistent in your own project (compact). Codex auto-loads AGENTS.md from
the working directory and ~/.codex/AGENTS.md globally. Copy this repo's
AGENTS.md into your project (or ~/.codex/AGENTS.md to load it everywhere) to
run routine J1-J6 review against your own codebase:
# per project
curl -o AGENTS.md https://raw.githubusercontent.com/vinicq/falsegreen-skill/master/AGENTS.md
# or globally, for every project
curl -o ~/.codex/AGENTS.md https://raw.githubusercontent.com/vinicq/falsegreen-skill/master/AGENTS.md
AGENTS.md carries the compact protocol and is self-sufficient for routine
review. Its on-demand references to reference.md and SKILL.md (the full
per-language catalog and long-form protocol) resolve only when those files sit
beside it, so for deep look-alike checks either drop reference.md/SKILL.md
next to the copied AGENTS.md or run from the clone above.
On the plugin marketplace. Codex does have plugin subcommands
(codex plugin marketplace add <source>, codex plugin add <plugin>@<marketplace>,
list/remove/upgrade, plus the /plugins TUI panel). This repo is itself
the plugin, so its catalog at .agents/plugins/marketplace.json uses a
repo-root source ("path": "./"). Codex's marketplace resolver expects a plugin
nested in a subdirectory of the marketplace root, so codex plugin marketplace add
does not install this repo as-is - use one of the two paths above. If you install it
from a conforming marketplace, qualify the plugin with the catalog name
(falsegreen): codex plugin add falsegreen-skill@falsegreen.
Note: if you were relying on a ~/.codex/prompts/ custom prompt, the plugin or
AGENTS.md path above is the maintained way to load this protocol. Keep using
your own prompts if they work for you; nothing here requires removing them.
Compact load order for Codex (context budget)
Codex loads project guidance into a host context with a working budget of
about 32 KiB. The full set of protocol files does not fit: SKILL.md
(~36 KiB) already breaks the budget on its own, and with AGENTS.md (~25 KiB)
plus this guide (~25 KiB) the three come to roughly 86 KiB loaded together.
Loading them at once truncates the protocol mid-file and the analysis degrades
silently.
Load the compact path instead. It carries the same J1-J6 protocol and case catalog through the single-source fragments, not a separate summary, so it cannot drift from the canonical text.
Eager (always loaded when Codex opens the project):
AGENTS.md- the project pointer. Codex reads it automatically at the repo root. It carries the compact J1-J6 protocol, the compact semantic-case table, and the precision-first rules. The semantic-case table and the precision rules are injected fromfragments/semantic-cases-compact.mdandfragments/precision-rules.mdbyscripts/sync-host-files.mjs, so they stay byte-identical to the canonical fragments. This file alone is enough to run the protocol on a typical file and fits the budget on its own.
On demand (load only when the case calls for it, never eagerly):
reference.md- the full per-language pattern catalog with examples and look-alike exemptions. At ~92 KiB it never fits eagerly, and neither does a whole language section:AGENTS.md(~25 KiB) plus the TS/JS section (~19 KiB) is ~44 KiB, past the budget before any test source loads, and Robot (~15 KiB) is over too. You do not need either.AGENTS.mdcarries the complete structural code index and the complete semantic table, so it names every code on its own; come here for the passage that defines a code only when a finding needs its full definition or an exemption the compact tables do not spell out. Two sections matter here and they are not interchangeable. The per-language section carries the structural codes for the file in front of you. The section## Patterns only the semantic pass can catch (AI-only)carries the language-agnostic S-series (S1-S18 and S21), which applies to every language including Python; the compact S-table inAGENTS.mdcovers it row-for-row, so pull the prose only when a finding needs the full definition or the "Look-alikes - do NOT flag" exemptions that close the section.SKILL.md- the full prose protocol, edge cases, and multi-agent mode. Load it only when you need the long-form judgment wording or the multi-agent procedure; the compact protocol inAGENTS.mdcovers routine review.
Single-source rule: the compact path MUST reference the J1-J6 protocol and the
case catalog through AGENTS.md (which is synced from the fragments/*
single sources) and through reference.md on demand. Do not fork a separate
compact summary of the protocol or the case table into this guide or anywhere
else - a forked summary drifts from the canonical text the moment either side
is edited. If a fragment changes, run npm run sync:hosts to re-inject it.
Model recommendations
Codex routes to OpenAI's current default model (the GPT-5 family at the time of writing). The CLI picks the version; this guide does not pin one, because a hard version string goes stale on the next release. The table below maps the three analysis passes to the capability you need, not to a frozen model id.
| Use case | Capability | models.yaml reference |
|---|---|---|
| Default (production review) | Codex's current default model | gpt-5 (semantic tier) |
| Fast / cheap batch | A smaller, faster sibling (a mini variant) | gpt-5-mini (structural tier) |
| Reasoning-heavy, case 18 analysis | A reasoning-tier model, extended reasoning on | gpt-5, reasoning tier (adversarial) |
The models.yaml column names the id validated for this release; if your
account exposes a newer default, prefer it - the capability, not the frozen id,
is what matters.
The current default handles all six judgments reliably, including the semantic cases (10, 11, 12, 15, 18). Drop to a smaller sibling when throughput or cost matters more than precision on edge cases. Use a reasoning-tier model when a case 18 finding needs extended chain-of-thought to cite an oracle and run an adversarial check.
For the API examples below, set model to the id your account exposes for the
current default (or its reasoning tier). See models.yaml for the canonical
tier-to-capability mapping the docs are validated against.
Note on reasoning models: some OpenAI reasoning models (the o-series) do
not accept a system message. When using one, fold the skill protocol into
the first user message instead. See the reasoning-model section below.
1. ChatGPT (chat.openai.com)
One-off review
- Open chat.openai.com.
- Paste the full contents of
SKILL.mdat the start of the conversation, followed by a blank line. - Paste the test file or snippet you want to analyze.
- Send.
The model will work through Steps 1-6 of the protocol and produce a report
in the CASE N (JX) - HIGH|LOW format with a SUMMARY block at the end.
Persistent context with Projects
ChatGPT Projects let you pin a system instruction that persists across all
conversations in that project. Use this to avoid pasting SKILL.md every
time.
- Create a new Project in ChatGPT.
- Open Project Instructions (gear icon or project settings).
- Paste the full text of
SKILL.mdinto the instructions field. - Save.
From that point on, every conversation in the project starts with the skill protocol loaded. You only need to paste the test code.
Tip: name the project something like falsegreen review so it is easy to
open when you want a quick analysis during a PR review.
2. OpenAI API (Python)
Basic usage
from pathlib import Path
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from environment
skill_protocol = Path("SKILL.md").read_text(encoding="utf-8")
test_code = Path("tests/test_example.py").read_text(encoding="utf-8")
response = client.chat.completions.create(
model="gpt-5", # Codex's current default model; use the id your account exposes
messages=[
{"role": "system", "content": skill_protocol},
{"role": "user", "content": test_code},
],
)
print(response.choices[0].message.content)
Reasoning models (o-series)
Some OpenAI reasoning models do not accept a system role. Combine the
protocol and the test code in a single user message:
from pathlib import Path
from openai import OpenAI
client = OpenAI()
skill_protocol = Path("SKILL.md").read_text(encoding="utf-8")
test_code = Path("tests/test_example.py").read_text(encoding="utf-8")
user_message = f"{skill_protocol}\n\n---\n\n{test_code}"
response = client.chat.completions.create(
model="o3", # any current OpenAI reasoning model that omits the system role
messages=[
{"role": "user", "content": user_message},
],
)
print(response.choices[0].message.content)
Use a reasoning model selectively for individual tests where case 18 is suspected, not for full-file batch runs - latency and cost are significantly higher.
3. Structured output
When you need machine-readable results — for CI integration, dashboards, or dataset collection — use OpenAI's structured output feature with a JSON schema.
Schema
{
"name": "falsegreen_report",
"schema": {
"type": "object",
"properties": {
"findings": {
"type": "array",
"items": {
"type": "object",
"properties": {
"case": { "type": "string" },
"judgment": { "type": "string", "enum": ["J1","J2","J3","J4","J5","J6"] },
"confidence": { "type": "string", "enum": ["HIGH","LOW"] },
"language": { "type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"] },
"level": { "type": "string", "enum": ["unit","integration","e2e"] },
"intent": { "type": "string", "enum": ["spec","char","regression","behavior"] },
"test": {
"type": "object",
"properties": { "name": { "type": "string" } },
"required": ["name"],
"additionalProperties": false
},
"finding": { "type": "string" },
"evidence": { "type": "array", "items": { "type": "string" } },
"oracle": { "type": "string" },
"fix_hint": { "type": "string" }
},
"required": ["case","judgment","confidence","language","level","intent","test","finding","evidence","fix_hint"],
"additionalProperties": false
}
},
"summary": {
"type": "object",
"properties": {
"tests_reviewed": { "type": "integer" },
"high": { "type": "integer" },
"low": { "type": "integer" },
"clean": { "type": "integer" }
},
"required": ["tests_reviewed","high","low","clean"],
"additionalProperties": false
},
"language": { "type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"] },
"framework": { "type": "string" }
},
"required": ["findings","summary","language","framework"],
"additionalProperties": false
},
"strict": true
}
Python usage
import json
from pathlib import Path
from openai import OpenAI
client = OpenAI()
skill_protocol = Path("SKILL.md").read_text(encoding="utf-8")
test_code = Path("tests/test_example.py").read_text(encoding="utf-8")
schema = {
"name": "falsegreen_report",
"schema": {
"type": "object",
"properties": {
"findings": {
"type": "array",
"items": {
"type": "object",
"properties": {
"case": {"type": "string"},
"judgment": {"type": "string", "enum": ["J1","J2","J3","J4","J5","J6"]},
"confidence": {"type": "string", "enum": ["HIGH","LOW"]},
"language": {"type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"]},
"level": {"type": "string", "enum": ["unit","integration","e2e"]},
"intent": {"type": "string", "enum": ["spec","char","regression","behavior"]},
"test": {
"type": "object",
"properties": {"name": {"type": "string"}},
"required": ["name"],
"additionalProperties": False,
},
"finding": {"type": "string"},
"evidence": {"type": "array", "items": {"type": "string"}},
"oracle": {"type": "string"},
"fix_hint": {"type": "string"},
},
"required": ["case","judgment","confidence","language","level","intent","test","finding","evidence","fix_hint"],
"additionalProperties": False,
},
},
"summary": {
"type": "object",
"properties": {
"tests_reviewed": {"type": "integer"},
"high": {"type": "integer"},
"low": {"type": "integer"},
"clean": {"type": "integer"},
},
"required": ["tests_reviewed","high","low","clean"],
"additionalProperties": False,
},
"language": {"type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"]},
"framework": {"type": "string"},
},
"required": ["findings","summary","language","framework"],
"additionalProperties": False,
},
"strict": True,
}
response = client.chat.completions.create(
model="gpt-5", # Codex's current default model; use the id your account exposes
messages=[
{"role": "system", "content": skill_protocol},
{"role": "user", "content": test_code},
],
response_format={"type": "json_schema", "json_schema": schema},
)
report = json.loads(response.choices[0].message.content)
for f in report["findings"]:
print(f"CASE {f['case']} ({f['judgment']}) - {f['confidence']} - {f['test']['name']}: {f['finding']}")
s = report["summary"]
print(f"\nSUMMARY: {s['tests_reviewed']} reviewed, {s['high']} high, {s['low']} low, {s['clean']} clean")
Schema field guide
| Field | Description |
|---|---|
case | Pattern code: C1, C3, 10, 18, etc. |
judgment | The first judgment that failed: J1 through J6 |
confidence | HIGH (no plausible legitimate interpretation) or LOW (likely smell) |
language | Python, TypeScript, or JavaScript |
intent | spec, char, regression, or behavior |
test.name | Name of the test function |
finding | One sentence describing what is wrong |
evidence | Array of specific line(s) that triggered the finding |
oracle | Required only for semantic case 18; not used for structural code C18 |
fix_hint | One sentence suggestion |
summary.tests_reviewed | Total number of test functions analyzed |
summary.high | Count of HIGH-confidence findings |
summary.low | Count of LOW-confidence findings |
summary.clean | Count of tests with no findings |
4. Codex CLI
If you installed the plugin, Codex discovers skills/falsegreen-skill/SKILL.md
through .codex-plugin/plugin.json. If you only cloned this repo, Codex reads
AGENTS.md as project guidance; that is useful, but it is not the same as an
installed skill. The options below cover running the protocol in a project that
has neither.
Per-session context
The Codex CLI has no --context flag. To load the protocol for a one-off run,
pipe the test file on stdin (-) and reference the skill from the prompt, or
prepend SKILL.md to the input:
codex "Apply the falsegreen-skill J1-J6 protocol from SKILL.md to the test file on stdin, and report false-positive smells" - < tests/test_example.py
To feed the protocol text inline, concatenate SKILL.md and the test file:
cat SKILL.md tests/test_example.py | codex "Analyze the test file below the protocol for false-positive smells" -
For a persistent setup, put the skill reference in AGENTS.md (next section) so
every session in the project picks it up without repeating it on the command line.
Project-level configuration
Add a section to your project's AGENTS.md (Codex CLI reads it automatically
when present) that points Codex to the skill:
## Test quality analysis
To analyze a test file for false-positive test smells, apply the
falsegreen-skill J1-J6 protocol from `SKILL.md`. Always follow
the six steps in order: detect language, apply Python catalog if Python,
classify test intent, apply J1-J6, adversarial-verify case 18, report.
Output findings as: CASE N (JX) - HIGH|LOW / Test / Finding / Evidence / Fix hint.
End with a SUMMARY block.
Then invoke:
codex "analyze tests/test_example.py for false-positive smells"
Codex will load the project AGENTS.md context and apply the protocol.
Test discovery
When the plugin is installed or AGENTS.md is present, Codex can find test
files automatically — you do not need to list paths. Say:
- "find and analyze all test files in this project"
- "run falsegreen on every test under tests/"
- "check the component tests in src/tests/"
Codex runs shell commands to discover files by pattern:
| Language | Patterns |
|---|---|
| Python | test_*.py, *_test.py |
| TypeScript / TSX | *.test.ts, *.spec.ts, *.test.tsx, *.spec.tsx |
| JavaScript / JSX | *.test.js, *.spec.js, *.test.jsx, *.spec.jsx |
Frontend component tests (React, Vue, Angular) match the same patterns and are analyzed with the same J1-J6 protocol as backend tests.
5. Batch processing
For large test suites, split by file and run API calls in parallel using
asyncio. This keeps total wall-clock time close to the slowest single file
rather than the sum of all files.
import asyncio
import json
from pathlib import Path
from openai import AsyncOpenAI
client = AsyncOpenAI()
skill_protocol = Path("SKILL.md").read_text(encoding="utf-8")
SCHEMA = {
"name": "falsegreen_report",
"schema": {
"type": "object",
"properties": {
"findings": {
"type": "array",
"items": {
"type": "object",
"properties": {
"case": {"type": "string"},
"judgment": {"type": "string", "enum": ["J1","J2","J3","J4","J5","J6"]},
"confidence": {"type": "string", "enum": ["HIGH","LOW"]},
"language": {"type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"]},
"level": {"type": "string", "enum": ["unit","integration","e2e"]},
"intent": {"type": "string", "enum": ["spec","char","regression","behavior"]},
"test": {
"type": "object",
"properties": {"name": {"type": "string"}},
"required": ["name"],
"additionalProperties": False,
},
"finding": {"type": "string"},
"evidence": {"type": "array", "items": {"type": "string"}},
"oracle": {"type": "string"},
"fix_hint": {"type": "string"},
},
"required": ["case","judgment","confidence","language","level","intent","test","finding","evidence","fix_hint"],
"additionalProperties": False,
},
},
"summary": {
"type": "object",
"properties": {
"tests_reviewed": {"type": "integer"},
"high": {"type": "integer"},
"low": {"type": "integer"},
"clean": {"type": "integer"},
},
"required": ["tests_reviewed","high","low","clean"],
"additionalProperties": False,
},
"language": {"type": "string", "enum": ["Python","TypeScript","JavaScript","Robot"]},
"framework": {"type": "string"},
},
"required": ["findings","summary","language","framework"],
"additionalProperties": False,
},
"strict": True,
}
async def analyze_file(path: Path) -> dict:
test_code = path.read_text(encoding="utf-8")
response = await client.chat.completions.create(
model="gpt-5-mini", # smaller sibling of the default; use the full default for higher precision
messages=[
{"role": "system", "content": skill_protocol},
{"role": "user", "content": test_code},
],
response_format={"type": "json_schema", "json_schema": SCHEMA},
)
report = json.loads(response.choices[0].message.content)
return {"file": str(path), **report}
async def analyze_suite(test_dir: str) -> list[dict]:
paths = list(Path(test_dir).rglob("test_*.py")) + \
list(Path(test_dir).rglob("*_test.py")) + \
list(Path(test_dir).rglob("*.test.ts")) + \
list(Path(test_dir).rglob("*.spec.ts"))
tasks = [analyze_file(p) for p in paths]
return await asyncio.gather(*tasks)
if __name__ == "__main__":
results = asyncio.run(analyze_suite("tests/"))
total_high = sum(r["summary"]["high"] for r in results)
total_low = sum(r["summary"]["low"] for r in results)
total_rev = sum(r["summary"]["tests_reviewed"] for r in results)
print(f"Files analyzed: {len(results)}")
print(f"Tests reviewed: {total_rev}")
print(f"HIGH findings: {total_high}")
print(f"LOW findings: {total_low}")
for result in results:
if result["summary"]["high"] > 0:
print(f"\n{result['file']}")
for f in result["findings"]:
if f["confidence"] == "HIGH":
print(f" CASE {f['case']} ({f['judgment']}) - {f['test']['name']}: {f['finding']}")
Practical notes:
- A smaller sibling of the default (a
minivariant) is the right default for batch runs. It is markedly cheaper than the full default and handles the structural families (A-E) accurately. Switch to the full default when reviewing files that are likely to contain semantic cases (10, 11, 12, 15, 18). - Split files larger than ~300 lines into logical groups before sending. The model's precision degrades when a single message contains too many test functions.
- Add a semaphore (
asyncio.Semaphore) to cap concurrent requests if you hit rate limits:sem = asyncio.Semaphore(10) # max 10 concurrent requests async def analyze_file(path: Path) -> dict: async with sem: ... # rest of the function unchanged
Related files
SKILL.md- the full J1-J6 protocol (system prompt)reference.md- per-language case catalogproviders.md- all supported LLM providerscontexts/- provider-specific context files