light-ocr

July 27, 2026 · View on GitHub

Core CI License C++17 Node--API v8 npm

arcships%2Flight-ocr | Trendshift

English | 简体中文

light-ocr pixel-art banner

Fast, offline OCR for Node.js and C++.

Recognize text in PDF, JPEG, PNG, or raw image data directly on your machine. light-ocr returns lines in reading order with confidence scores and quadrilateral coordinates. For Node.js, the npm package includes PP-OCRv6 Small, PDFium, and prebuilt components for macOS, Linux, and Windows.

Quick start

Node.js 22 and 24 are supported.

npm install @arcships/light-ocr
import { createEngine } from "@arcships/light-ocr";
import { readFile } from "node:fs/promises";

const engine = await createEngine();
const result = await engine.recognizeEncoded(
  await readFile("image.jpg"),
);

for (const line of result.lines) {
  console.log(line.text, line.confidence, line.box);
}

await engine.close();

createEngine() automatically chooses the right execution mode for the current platform. If your application already decodes images, recognize() also accepts GRAY8, RGB8, BGR8, and RGBA8 pixel data.

Model tiers

Small remains the stable default. N2 also provides two opt-in preview packages under the next tag; all three expose the same API, types, result schema, and error model, while each install contains only its selected model.

TierPackage / commandModel payloadStatus
Small@arcships/light-ocr / light-ocr~30 MBstable default
Tiny@arcships/light-ocr-tiny@next / light-ocr-tiny~6.3 MBpreview; 49 languages, no Japanese
Medium@arcships/light-ocr-medium@next / light-ocr-medium~139 MBpreview; quality-first

Tiny and Medium stay on next until real use shows a clear reason to promote them; they do not change what npm install @arcships/light-ocr installs.

CLI

The light-ocr command is available after install — no extra setup:

# Recognize text + coordinates
light-ocr image.png --format json

# Just text
light-ocr image.png --format text

# PDF pages, using the renderer already included by npm
light-ocr report.pdf --pages 1-10 --format text

# Detect text regions only (no recognition)
light-ocr detect image.png

# Region of interest
light-ocr recognize image.png --region 100,80,640,320 --format json

# Engine info
light-ocr info --version

# System diagnostics (hardware and providers)
light-ocr doctor --json

Image commands are recognize (default), detect (boxes only), info (version diagnostics), and doctor (system diagnostics). A .pdf path routes directly to document OCR; document handles explicit multi-source jobs. Output uses a versioned schemaVersion: 1 contract. EXIF orientation is corrected automatically. See the CLI design and npm README for full reference.

PDF and multi-page documents

PDF and multi-page OCR are built into @arcships/light-ocr. The matching PDFium binary is carried by the same platform npm package as the OCR runtime: there is no postinstall script, runtime download, compiler, or separate package to install.

# Single PDF with default 150 DPI
light-ocr report.pdf

# Page range with streaming JSONL output
light-ocr report.pdf --pages 1-10 --format jsonl

# Multiple images as one document
light-ocr document scan1.png scan2.png scan3.png --format text

Programmatic API:

import { recognizeDocument } from "@arcships/light-ocr";

// Stream pages from a PDF
for await (const page of recognizeDocument("report.pdf", { dpi: 200 })) {
  console.log(page.index, page.lines.length, page.source.kind);
}

// Multiple images
for await (const page of recognizeDocument([buf1, buf2, buf3])) {
  console.log(page.index, page.lines);
}

Agent Skill

An Agent Skill is included for AI agents that can call local commands. It provides scenario-driven workflows, a decision flow for command selection, and exit code reference:

  • When to use OCR vs. a multimodal model
  • Detect-then-recognize two-step pattern for large images
  • ROI field extraction for receipts and forms
  • Verifying multimodal output against deterministic OCR

What you get

  • Local processing. Images, PDFs, and OCR results stay on your machine.
  • One package to install. The model, OCR runtime, and PDF renderer are included through the npm package's platform dependency.
  • No secondary downloads. Installation and runtime need no postinstall fetch, compiler, model download, or PDF engine download.
  • Useful output. Every line includes recognized text, confidence, and its position in the original image.
  • Hardware acceleration by default. Auto tries Core ML first on macOS 15+ Apple Silicon, and WebGPU first on the Linux and Windows builds below.
  • Application-friendly execution. Recognition runs off the JavaScript main thread and supports queues, cancellation, and explicit cleanup.
  • Small text in large images. An optional tiled mode preserves small and dense text in high-resolution images.

Like light-ocr? Give it a star — it helps others discover the project and keeps us motivated!

Platform acceleration

The npm package provides the following six builds. The default createEngine() call uses Auto mode:

PlatformAuto mode
macOS on Apple SiliconCore ML on macOS 15+, then CPU
macOS on IntelCPU
Linux x64 with glibcWebGPU through Vulkan, then CPU
Linux arm64 with glibcCPU
Windows x64WebGPU through D3D12, then CPU
Windows arm64CPU

Applications that need explicit control can choose auto, cpu, apple, or webgpu through the execution option.

Measured performance

Version 0.3.0 was measured on three real devices:

light-ocr 0.3.0 same-device speed and OCR process CPU-time reductions

DeviceAccelerationEnd-to-end speedupOCR process CPU time
Apple M4 MaxCore ML2.30× on HELLO 123; 2.85× on a dense form95.91%–97.67% less
NVIDIA RTX 5060 Ti on LinuxWebGPU / Vulkan5.70× overall across 14 test images69.97% less
AMD Radeon 780M on WindowsWebGPU / D3D122.44× overall across 14 test images46.33% less

These are same-machine comparisons with the CPU path, and results vary by workload and hardware. For the 14-image results, overall speedup is the sum of the per-image CPU median times divided by the sum of the WebGPU median times. The CPU column measures cumulative OCR process CPU time over the same workloads, rather than an instantaneous system-utilization sample; lower CPU time leaves more capacity for the rest of the application while OCR is active. The Apple run passed its locked CPU-parity thresholds; both WebGPU runs were byte-identical to CPU FP32 on all 14 images. See the 0.3.0 release report for complete measurements and methodology.

C++

C++ projects build the static library from source and link the light_ocr::core CMake target. The API accepts decoded GRAY8, RGB8, BGR8, or RGBA8 pixels; start with the C++ API guide and build instructions.

Documentation

Community and license

Issues and pull requests are welcome — see CONTRIBUTING.md for guidelines. All participants are expected to follow our Code of Conduct.

light-ocr is available under the Apache License 2.0.