jgrep

September 20, 2026 ยท View on GitHub

grep, but the pattern is a description.

$ tail -f app.log | jgrep "a user is getting frustrated"
user 12: this is the third time checkout has failed, I am done with this app
user 77: WHY does it log me out every five minutes??

$ jgrep -o "announces or releases a new AI model" titles.txt | sort -rn | head -3
0.980	PrismML Launches Bonsai 2 27B, Its Most Capable Model Yet
0.970	Alibaba Releases Qwen3.8-Omni-Flash
0.940	Google announces new experimental "CC" AI agent for families

Each line becomes one yes/no question to Jev, TypeSafe's decision model. Jev does not generate text. It returns a probability in about 200 ms for about a thousandth of a cent, which is fast and cheap enough to sit in a pipe. jgrep reads lines as they arrive, judges them concurrently and prints matches in input order, so it works on tail -f as well as on files.

Measured on 994 Hacker News titles: 4.6 seconds and $0.012 for one description, and the same time for three descriptions at once.

Install

uv tool install jev-grep        # the command it installs is jgrep
uv tool upgrade jev-grep        # upgrade an existing installation

For Go function parsing, install the optional syntax parser: uv tool install 'jev-grep[code]'. Python function parsing and unified diffs work with the base package.

jgrep needs a key for one of two APIs, or for a gateway of your own (below). With keys for several, it uses TypeSafe's.

APIKeyGet one
TypeSafeTYPESAFE_API_KEYconsole.typesafe.ai
OpenRouterOPENROUTER_API_KEYopenrouter.ai/keys

Set the environment variable, or put the key in ~/.config/jev/typesafe.key or ~/.config/jev/openrouter.key. Force a choice with --api or JEV_API.

Behind an LLM gateway that serves System One (LiteLLM, Ramp Router, a corporate proxy), point jgrep at it with --api gateway. The URL is the full endpoint and the key is the gateway's own:

export JEV_GATEWAY_URL=https://gateway.example.com/v1/systemone
export JEV_GATEWAY_API_KEY=...       # or ~/.config/jev/gateway.key and gateway.url
jgrep "a stack trace" build.log      # picked automatically when it is the only key set

Requests are sent as they would be to TypeSafe, so the gateway sees the same {model, state, questions} body. Ask for a model the gateway knows with --model; the default is jev-latest. --stats prices gateway calls at TypeSafe's list price, which may not be what the gateway bills.

Use

jgrep "a complaint about noise" complaints.txt          # lines that fit
jgrep -v "spam" inbox.txt                               # lines that do not
jgrep -c "asks a question" *.txt                        # counts per file
jgrep -p 0.9 "mentions a specific dollar amount" f.txt  # only confident matches
jgrep -o -p 0 "the writer is losing sleep" f.txt | sort -rn   # rank every line
jgrep -e "about economics" -e "about New York" f.txt    # either; add --all for both
jgrep -C 2 "a line in the middle of a stack trace" app.log   # judged with its neighbours
jgrep --para "describes an identification strategy" paper.txt
jgrep --whole "uses a bunching estimator" abstracts/*.txt     # prints matching file names
jgrep -q "a stack trace" build.log && notify "build broke"
jgrep --jsonl --field message "a payment failed" events.jsonl
jgrep --csv --field abstract "uses a natural experiment" papers.csv
jgrep -rl --glob '*.txt' "mentions a rent increase" notes/
jgrep --chunks 8000 --json "describes an identification strategy" paper.txt
OptionMeaning
-p PMatch when the probability is at least P. Default 0.5.
-oPut the probability in a first, tab-separated column.
-v, -c, -n, -H, -qAs in grep.
-m NUMStop each input file after NUM matches; -m 0 reads no input and makes no API calls.
-lPrint a file name at its first matching record, then continue to the next file.
-rSearch directories recursively; defaults to the current directory if no paths are given.
--glob PATTERN, --exclude PATTERNInclude or exclude files; both can be repeated.
--no-ignoreDuring recursive search, disregard .gitignore and .ignore.
-e DESCAnother description. All of them go in one call per line. A line matches if any fits, or all with --all.
--para, --wholeJudge paragraphs or whole files in place of lines.
-C NShow Jev the N lines either side of each line. Still one decision per line, and still only the matching line is printed.
--jsonOne JSON object per match, with the probability.
--jsonl --field NAME, --csv --field NAMEJudge one field and return the complete original record.
--chunks N, --overlap NSearch full text files in overlapping passages, with source locations.
--diffJudge each complete unified diff hunk, including removed lines and unchanged context.
--functions, --lang python|goJudge complete functions/methods with adjacent comments; infer language from the extension, or specify it for stdin.
--estimateRead to EOF and preview calls and approximate cost without authentication or API calls. Add --json for a single report.
--emit-recordsExport source-linked JSONL without judging; omit DESCRIPTION. Useful for inspection and other tools.
--max-chars NMaximum characters judged per ordinary record; default 8000. Truncation produces a warning.
--unorderedPrint matches as answers arrive.
-j NCalls in flight. Default 32.
--budget DOLLARSStop once this much is spent. Default 1.00, or $JGREP_BUDGET; 0 for no limit.
--timeout SECONDSGive up on a line after this long, retries included. Default 15.
--no-cache, --api, --model, --statsSee jgrep --help.

Exit status follows grep: 0 if anything matched, 1 if nothing did, 2 on error. Offline estimation and record export exit 0 on success, including empty input, and 2 on errors.

With ordered output, -j N bounds the total of active requests and completed results waiting for earlier records. A slow first record therefore cannot let the rest of the input run ahead. --unordered releases each slot as soon as its result arrives.

Changes and complete functions

# Review both sides of each change, including deletion-only hunks
git diff --no-color | jgrep --diff "removes error handling for a persistent write"
jgrep --diff "weakens cancellation handling" review.patch --json

# Read the entire function; comments and decorators stay attached
jgrep --functions "ignores a failed rollback" plugin/installer.go --json
jgrep --functions "releases connections on every exit path" src/ -r --glob '*.py'

# Preview exactly the units the filter would judge; no key or paid call needed
git diff --no-color | jgrep --diff --estimate "removes error handling" --json

# Export functions for jselect without making any jgrep model calls
jgrep --functions --emit-records src/ -r > functions.jsonl
jselect "How does cancellation work?" functions.jsonl --tokens 2000

--diff accepts ordinary unified patches, including Git diffs, from files or stdin. One decision covers a complete hunk: both removed and added lines, plus the context supplied in the patch. Use git diff -U10 when you need more surrounding lines. Each emitted hunk repeats its file headers. It does not read the working tree or retrieve omitted context. Binary changes, metadata-only changes (such as mode-only edits), combined merge diffs, and malformed hunks produce errors rather than silently reporting no match. An empty diff contains no records.

For diff JSON, file, line, and end_line locate the hunk in the input patch, while unit contains old_file, new_file, old_start, old_count, new_start, and new_count. Missing file sides are null; zero-length ranges retain the unified diff's insertion/deletion anchor. -c counts matching hunks per input patch and -l names matching input patches.

--functions supports Python through the standard-library AST and Go through the optional Tree-sitter parser. It extracts named functions and methods, retaining decorators, docstrings, adjacent comments, and nested function bodies. Nested functions are not emitted again separately. Imports, class-level state and callers are not automatically attached. Recursive discovery skips other extensions unless --lang is explicit. Syntax errors are reported rather than guessed around. Function JSON includes unit.language, unit.symbol, and exact decoded-character start/end offsets with one-based source lines. The Python API exposes the same deterministic readers in jgrep.code_inputs.function_records and diff_records.

Diffs and functions never truncate: units over --max-chars fail with their size and location. Raise that limit deliberately if needed. These modes cannot combine with -C, --para, --whole, --chunks, or structured input. They read each input file into memory; keep live streams in line mode.

--emit-records writes one object containing schema_version, id, text, source, line/span locations and unit. IDs include the source location and a text hash, so tools that select only text and id, including jselect, retain a traceable reference. It exports full parsed records, without ordinary line-mode truncation or a relevance judgment. Do not supply a description or match filters. Errors are JSONL objects with an error.message and cause exit 2; a pipeline must check the producer's exit status before treating its export as complete.

These modes retrieve evidence for review. Model scores do not prove a bug or certify that a change is safe. Results and observed failures on 20 handwritten examples are in the code-review experiment.

Cost preview

--estimate uses the same input mode, selected field, context window, descriptions and model as the filter. It reads existing cached answers in read-only mode and estimates reuse of exact repeated requests. It does not normalize whitespace or identifiers, create a cache, or contact the provider. Without a configured provider it uses TypeSafe's default model; use --api and --model to preview a specific setup. No API key is required.

The JSON report includes record/cache/duplicate counts, estimated_calls, call_upper_bound before new duplicate reuse, estimated_input_tokens, estimated_cost_usd, byte_estimate_cost_usd, errors and assumptions. Costs use JEV_PRICE_PER_MTOK (default $0.042 per million input tokens). The nominal estimate is UTF-8 request bytes divided by four plus 270 overhead tokens per request; the broader byte estimate uses those bytes plus 1,024 overhead. Neither is a provider quote or guaranteed cap. Retries, gateway pricing and concurrent cache misses can change actual cost.

Preview reads to EOF and ignores matching stop conditions (-q, -l, -m, threshold and budget), because their effects depend on model answers. It reports ordinary-record truncation and oversized code-unit errors. Use finite input, not an endless tail -f stream. With --json, the preview is one object; errors makes a partially readable collection explicit and the exit status is 2.

Structured records

--jsonl expects one JSON object per line. --csv expects a header with unique column names and supports quoted commas and multiline cells. Both require --field NAME. Only that field is sent for judgment; -C also supplies that field from neighboring records. Matching output retains the complete JSON line or CSV row, including fields that were not judged. CSV output includes the original header once for each input file with matches. Record terminators are written as newlines, while quoting and embedded newlines are preserved.

JSON field names can be dotted paths, such as event.message or events.0.message. An exact key takes precedence over a dotted path. Strings are judged directly, null is treated as empty text, and other values are represented as JSON. Missing fields, invalid JSON, and malformed CSV rows are reported as errors; valid later records are still processed when parsing can continue. Blank JSONL lines are skipped.

Structured output omits automatic filename prefixes so it can be read as JSONL or CSV. Use -H or -n only when you want those prefixes. --json instead emits a match object containing file, the starting physical line, p, text (the original record), field, and record (the parsed object, with CSV column values kept as strings). For several CSV files with different headers, this JSON output is easier to combine. -c counts matching records, and -l lists matching files without headers or rows.

-r visits files in sorted order, respecting .gitignore and .ignore in each directory and parent ignore files inside a Git repository. It skips symlinks, .git directories, and files with a NUL byte in the initial binary check. --glob '*.txt' selects names anywhere below the search root; patterns can also match relative paths. --exclude uses gitignore patterns. --no-ignore disables ignore files but still honors explicit exclusions. Explicit file arguments bypass ignore files; include/exclude filters still apply. Use -l when only the matching paths are needed.

Ordinary lines, paragraphs, selected fields, and --whole judgments are limited to the first 8000 characters by default. jgrep reports when this truncates a record or its context. Raise --max-chars to change that limit, or use --chunks N to search the entire text file a passage at a time. Chunk size replaces the ordinary character limit. Overlap defaults to the smaller of 200 characters and one quarter of the chunk size; --overlap 0 disables it.

Chunk output includes the starting line number. With --json, each matching passage also has chunk (one-based), start and end (zero-based character offsets, end exclusive), and end_line (the last source line containing characters from the passage). Offsets count decoded Unicode characters, not bytes. Passages can overlap and are judged independently; their scores are not combined into a document-wide probability. -c counts matching passages, while --chunks 8000 -l lists files with at least one matching passage.

Chunk mode applies to plain text files and cannot be combined with structured input, --para, --whole, or -C; use overlap to retain text across passage boundaries. These modes read text, not PDF or Word formats.

Cost

A call bills roughly 270 tokens of fixed overhead plus the line and the description, so a typical line costs about 300 tokens, or $0.0000126 at $0.042 per million. A million lines is about $13. Blank lines, repeated lines and anything answered before are free: answers are cached in ~/.cache/jev/answers.sqlite, keyed on the provider, endpoint, exact model, judged text and description. Changing gateways cannot reuse another endpoint's answers. Older cache entries without provider/endpoint identity are not reused, so the first rerun may make fresh calls. Extra -e descriptions add about 27 tokens each and no time. -C N sends 2N+1 lines in place of one, so -C 2 costs roughly three times as much per line once the fixed overhead is counted.

jgrep stops at --budget, one dollar by default, so a stray jgrep pattern huge.log cannot run up a bill. A dollar is about 80,000 lines. A stopped run loses nothing: rerun with a higher budget and everything already judged comes from the cache. For a long-lived tail -f monitor, set your own default once with export JGREP_BUDGET=20, or 0 for no limit. With --stats, or whenever stderr is a terminal, it prints what the run cost:

jgrep: 994 records, 33 matched; 994 calls, 0 cached; 292,839 tokens; \$0.0123; 4.6s

How well does it work

Three benchmarks on public labeled text, run on 2026-09-18 with Jev 1.13 through OpenRouter. Each one runs the installed jgrep command itself, uncached, at its default threshold of 0.5. Reproduce them with bench/accuracy.py.

Against a keyword grep. The UCI SMS Spam Collection: 5,574 text messages, 747 of them spam.

FilterPrecisionRecallF1TimeCost
jgrep "an unsolicited spam, scam or marketing text message"0.870.950.9127 s$0.07
the same with -p 0.90.980.840.90
grep -iE "free|win|prize|claim|urgent|cash|txt|call now|..." (17 terms)0.640.810.720.03 sfree

The regular expression was written before looking at any results and is in the script.

Against asking a chat model. The do-it-yourself alternative is a loop that asks an LLM the same yes/no question about each line. On 300 of those messages, 32 requests in flight, all through OpenRouter:

JudgeF1Wall timeCostMedian latency
jgrep (Jev 1.13)0.902.7 s$0.0039about 210 ms
GPT Luna0.889.3 s$0.0060802 ms
GPT Terra0.9210.8 s$0.0571988 ms
Qwen 3.7 Flash, thinking off0.788.4 s$0.0007789 ms

jgrep finished three to four times sooner than any of them. Its accuracy sits between the two GPT tiers; with 45 spam messages in the sample, those three F1 scores are within noise of each other. It is not the cheapest per line: a small open model costs a sixth as much and is clearly less accurate. Against the model that matched its accuracy, jgrep cost a fifteenth as much.

Several descriptions at once. AG News test set, 7,600 articles, four descriptions (-e "news about sports" -e "news about business, markets or the economy" ...) judged in one call per article: 37 seconds and $0.13 for all four. Taking the most probable description as the label gives 86.6% accuracy with no training. One-vs-rest F1 at 0.5 was 0.97 for sports, 0.82 for science and technology, 0.82 for world affairs and 0.72 for business, which over-triggers (precision 0.58) because so much technology news is also business news.

Does the wording of a description matter? bench/phrasing.py scores 30 hand-labeled lines against five descriptions of different grammatical shapes, including a negation and a question. Jev got all 150 right under each of four ways of wording the question; that set is easy on purpose. Asking five descriptions in one call changed no decision and moved probabilities by 0.001 on average. Latency was flat at about 210 ms from 1 to 64 questions per call.

On borderline lines the probabilities land in between, which is what -p is for:

0.65  [a complaint about noise]  The music from the church on Sunday mornings is lovely but it does start early.
0.46  [does not mention a landlord]  The owner of the building never answers the phone.

Things to know:

  • These are a model's judgments. Check a sample before you rely on a filter.
  • Jev answers the description you wrote, not the one you meant. TypeSafe documents weak spots: counting, comparing numbers or dates, double negatives, and long inputs full of irrelevant detail.
  • A line is judged alone unless you pass -C N, which shows Jev the N lines either side of it. The decision is still per line, and the neighbours are not printed. Context does not reach across files, and on tail -f a line cannot be judged until the N after it arrive.
  • Jev is close to deterministic, not exactly so. Asking 150 questions three times without the cache gave identical probabilities for 128; the rest moved by up to 0.03 and no decision flipped. The cache makes reruns exact.
  • The default model ID is an alias for the latest Jev. For results that must reproduce, pin one with --model (for example typesafe/jev-1.13 on OpenRouter).
  • Text in the input can try to steer the answer. Do not use jgrep as a security boundary.

Development

uv sync && uv run pytest        # offline tests using a fake API and local HTTP server; no key
uv run python bench/phrasing.py # live; costs about a cent
uv run python bench/accuracy.py prepare && uv run python bench/accuracy.py spam   # also: news, llm

src/jgrep/inputs.py handles file discovery, structured records, and document passages. src/jgrep/core.py is the client: API backends, retries inside a time budget, the cache, in-flight deduplication and the cost meter. It shares its design with jlink, which links records across datasets with the same model.

MIT license.