DocJev

September 21, 2026 · View on GitHub

All projects · Developer tools

Classify or split PDF/DOCX/PPTX packets with local LiteParse OCR text and TypeSafe Jev category/boundary judgments; optional LlamaParse for hard scans.

At a glanceDetails
SourceSource
Maintainerjerryjliu. Independently curated; this page is not an upstream submission or endorsement.
FormatPython package/CLI docjev (alias jev-docs); local visual demo optional.
RequirementsPython ≥ 3.11, TYPESAFE_API_KEY. Optional LLAMA_CLOUD_API_KEY / OPENAI_API_KEY. DOCX/PPTX need LibreOffice.
LicenseApache-2.0. TypeSafe and optional cloud OCR have separate costs.
DisclosureAI-assisted catalog review; no affiliation. Listing is not an endorsement. Offline pytest run; live Jev classify/split and billed OCR not run.

When to use

Use it for inbox-style document classification and multi-document PDF splitting with typed Jev answers over extracted page text. Prefer doc-router when you only need page-level OCR routing rather than category/boundary judgments. Distinct from LlamaIndex hosted Classify/Split APIs (explicitly not used).

How it works

LiteParse (or optional LlamaParse) extracts page text; engines/jev.py calls TypeSafe via typesafe-sdk with Choice/Noul questions for category or split boundaries. Application code owns rules YAML, PDF export, and metrics. Jev inference is hosted even when OCR is local.

Get started

git clone https://github.com/jerryjliu/docjev.git
cd docjev
git checkout 9ed0fe05984ce1906af9272b8b400c8d46520f98
uv sync
# export TYPESAFE_API_KEY=...
uv run docjev doctor --smoke

Live classify/split examples are in the upstream README (incurs TypeSafe charges).

Examples and demos

  • Upstream README Quick start, rules YAML, and visual report under docs/report/.
  • Local demo: uv sync --extra demo then uv run docjev demo (needs keys for live runs).
  • This listing ran uv run pytest -q -m 'not live': 174 passed, 3 failed (tests/test_ocr_contract.py cloud-contract cases), 3 deselected. No live TypeSafe call.

Limits and data handling

Document text and rules reach TypeSafe for judgments. Optional LlamaParse/OpenAI paths send content to those providers. Treat upstream accuracy/timing tables as reported, not re-measured here.

Review and maintenance

Reviewed on 2026-09-21 at commit 9ed0fe0: Apache-2.0; AI-assisted source review of README, LICENSE, engines/jev.py, CLI; offline pytest as above. No live TypeSafe call.

Related: doc-router, jev-table, llama-index-jev.