Changelog
September 5, 2026 · View on GitHub
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Fixed
- Print only the exact pinned
pipx installcommand when the plugin launcher cannot execute DocPull.
[6.5.5] - 2026-09-04
Fixed
- Report the DocPull package version during MCP initialization and attach the
advertised cost description and
docpull/costmetadata to every MCP tool.
[6.5.4] - 2026-09-04
Fixed
- Treat a non-executable
docpullpath as an unavailable runtime and return the same actionable setup guidance as a missing executable.
[6.5.3] - 2026-09-04
Fixed
- Shorten the Codex activation prompt to the host's 128-character limit so the focused research trigger is loaded instead of ignored.
[6.5.2] - 2026-09-04
Fixed
- Add a dependency-free plugin launcher that reports the exact
pipxsetup command when thedocpullexecutable is unavailable.
[6.5.1] - 2026-09-04
Fixed
- Constrain the optional MCP runtime to the compatible 1.x series. This prevents
fresh
docpull[mcp]installations from resolving MCP 2.x, whose low-level server API is incompatible with DocPull 6.5.
[6.4.0] - 2026-07-16
Added
- Add
website-packacross CLI, Python workflow API, and MCP. The workflow emits the pinnedwebsite.snapshot.v1schema, stable URL-based document identities, content-hash versions, page roles, source authority, OKF/raw representations, bounded optional visuals, portable-v3 manifests, and recursively verifiable artifact hashes. - Add verified-baseline diff states for added, changed, unchanged, removed, failed, and blocked pages, plus golden/tamper fixtures.
Changed
- Let brand, product, policy, styleguide, image, and relationship workflows consume existing packs without network fallback. Product extraction now keeps trial metadata separate from ordinary price fields and excludes testimonial copy from feature evidence.
Removed
- Remove the standalone DocPull website, its Next.js workspace, and web-only CI/release gates. GitHub, PyPI, and the MCP Registry remain the canonical public product surfaces.
[6.3.0] - 2026-07-16
Added
- Add
relationship-packacross CLI, SDK, MCP, workflow registry, and project sources. It emits citedowned_by,operated_by,acquired_by,franchised_by, andinvested_inobservations plus one coverage result per input; missing evidence remains a coverage gap and never becomes an independence claim. - Put core
fetch,crawl, anddataset-packonWorkflowRequest/WorkflowResult; empty and partial crawls now retain a structured current-run manifest, progress, budget usage, hashes, replay configuration, warnings, and typed failures. - Support bounded remote HTTPS JSON/CSV dataset snapshots with original and resolved URL provenance, query parameters, content type, and deterministic SHA-256 hashes.
- Add the DocPull product site with home, pricing, privacy, terms,
llms.txt, robots, sitemap, manifest, branded icons, and reusable launch assets. - Add metadata-only native integration context adapters that produce cited, rights-labeled document records without storing provider secrets or raw customer payloads.
- Add a committed
uv.lock, pinned Mise runtimes, Next.js ESLint gates, and dependency review for reproducible local and CI environments.
Changed
- Make CLI success strictly current-run scoped by default. Add explicit
--exit-policy usable-outputfor consumers that intentionally accept older records in a shared directory. - Classify authority per entity/source in multi-company bundles and add
official_corporate,government_registry,regulatory_filing,press_release, andlocal_reportingroles. - Extract relationship observations into intelligence bundles while keeping every candidate unresolved until downstream human review.
- Load the root SDK and internal package exports lazily while preserving every documented import and CLI, Python SDK, and MCP surface.
- Coalesce byte-equivalent concurrent HTTP GETs without retaining completed responses, and isolate requests with different timeouts or headers.
- Journal frontier transitions between compact snapshots so interrupted runs remain resumable without rewriting the full frontier after every URL.
- Reuse pack reads, citation analysis, entity extraction, and graph indexes across intelligence workflows; use constant-time document/source lookups and bounded top-result selection for local search.
- Centralize schema versions and export-format registries in lightweight modules to reduce import cost and duplicated contract definitions.
Fixed
- Populate retryable acquisition failures with stable codes such as
http_429, fetch stage, HTTP status, attempt count, and Retry-After seconds. - Prevent stale records in a reused output directory from turning a zero-record current run into exit status 0.
- Restore automatic CI, CodeQL, security, benchmark, live-smoke, and metrics triggers after the runtime-tooling migration, and make the Python matrix select the advertised patch versions explicitly.
- Extend release-readiness checks to lint, type-check, audit, and build the web workspace, and repair stale repository links in issue templates.
Compatibility
- Keep all 6.2 pack, workflow, CLI, SDK, MCP, schema, and tracker contracts readable and importable; 6.3 changes are additive or internal optimizations.
[6.2.0] - 2026-07-16
Added
- Define DocPull as the evidence and acquisition engine and publish versioned
WorkflowRequest,WorkflowResult,ArtifactManifest,intelligence.bundle.v1, andChangeEventcontracts with bundled JSON Schemas. - Add generic progress, warnings, failures, budget usage, SHA-256 artifact hashes, pack/run identity, and scheduler-neutral replay settings.
- Promote brand, product, styleguide, image, screenshot, and new policy packs to public CLI, SDK, and MCP workflows, plus declarative project source types for brand, product, styleguide, visual, and policy evidence.
- Add policy discovery/classification, effective dates, stable clause hashes, and clause-level before/after candidates without legal conclusions.
- Add deterministic tracker bundles with source snapshots, document versions, precise evidence spans, observations, confidence, evidence strength, source authority, warnings, and change candidates.
- Add idempotent change events separating structural, textual, and semantic candidates for pricing, positioning, product, security, and policy changes.
Changed
- Enrich product pricing extraction with plans, currencies, billing intervals, trials, feature gating, page-text provenance, and incidental/visual price exclusion.
- Make JSON-only eval preparation retain the required
PACK_CARD.mdartifact. - Keep
company_brain.bundle.jsonand its SDK builder as compatibility aliases forintelligence.bundle.v1.
Compatibility
- Validate old packs remain readable and retain legacy result filenames while new workflow sidecars and schema fields are additive.
[6.1.0] - 2026-07-14
Added
- Add local
pypdfdocument extraction through thepdfextra, including isolated subprocess execution, resource and output limits, secure temporary files, structural PDF classification, and extraction-quality provenance. - Add scorer v4, portable report schema v3, explicit comparison scopes, encrypted evidence escrow, signed publication verification, and clean-wheel subject provenance to the isolated benchmark lab.
Changed
- Prefer
pypdffor remote PDF extraction before optional MarkItDown and Unstructured fallbacks, with consistent CLI, Python, and MCP controls. - Fence structured raw responses with adaptive labeled Markdown fences while preserving Markdown and documentation-style plain text directly.
- Make repository-hosted manual benchmark runs exploratory only; comparative claim evidence must come from the external sealed-holdout process.
Fixed
- Reject fused words in benchmark evidence while retaining explicit line-break hyphen repair and punctuation-separated token equivalence.
- Recompute benchmark summaries and billing from observations, rejecting forged, incomplete, duplicate, or tampered report and publication data.
[6.0.0] - 2026-07-01
Added
- Add DocPull output contract v3 validation with
docpull pack validate, raw/agent/eval levels, v3 manifests, raw sidecars for file-backed outputs, record-level citations, chunk-safe SQLite output, and listing item sidecars for link-dense event/news pages. - Add
docpull parse, a local document parsing lane that emits v3 packs from plain text/Markdown directly and optional MarkItDown or Unstructured parser backends for complex office/PDF files. - Add explicit
--remote-documents pdfsupport for locally parsing fetched PDFs while keeping remote documents, browsers, and cloud parsing disabled by default. - Add optional Presidio-backed PII detection for
docpull pack audit --redactionanddocpull pack redact, while keeping deterministic regex redaction as the default backend. - Add
docpull openapi-packto convert local or HTTPS OpenAPI JSON/YAML specs into v3 endpoint and component-schema records. - Add
docpull feed-packto convert RSS, Atom, JSON Feed, or pages that advertise feeds into item-level v3 records with feed item sidecars, listing sidecars, and freshness metadata. - Add typed knowledge lanes for known-source context dependencies:
docpull paper-pack,repo-pack,package-pack,standards-pack,dataset-pack,transcript-pack, andwiki-pack. These commands emit v3 raw packs, validate withdocpull pack validate, prepare to agent/eval grade, and export through the existing pack export formats without adding new MCP tools. Remote typed lanes also support opt-in metadata caching, typed sidecar roots, async SDK wrappers, standards section records, exact streamable dataset row counts, and explicit official API source contracts for arXiv, Crossref, and NCBI E-utilities. - Add
--extractor ensemble, which scores built-in and optional trafilatura extraction candidates and keeps the strongest Markdown output. - Add
docpull ci, a local Context CI gate for project and standalone pack context. It validates lockfiles, current score/audit sidecars, coverage, citation coverage, eval-grade artifacts, rights constraints, and optional context prediction pass rates while writingcontext-ci.report.jsonandCONTEXT_CI.md. - Add GitHub Actions workflow definitions for scheduled free live smokes over typed official APIs and ordinary web sources, so source drift can be checked outside the default offline test suite.
- Add
scripts/real_feature_smoke.py, an opt-in real-data acceptance harness for the free/local public surface with optional cloud-render lanes. - Document the Context CI GitHub Actions workflow, context-pack contract, and hidden-eval human review protocol.
- Add Context CI examples, Context Pack Contract v3, a design-partner playbook, and a spec-only hosted control-plane extension.
- Make
current-context-qathe default evalgen task type while acceptingcurrent-docs-qaas a legacy alias. - Add a Context CI benchmark report over local demo packs showing passing, manifest-mismatch, coverage, rights, citation, and stale-answer gates.
Changed
- Consolidate the unreleased v3 surface around context dependencies, pack
validation/preparation, exports, Context CI, document parsing, OpenAPI packs,
and the canonical
agent-browserrenderer contract. - Remove unreleased top-level parity/provider/benchmark commands from the public CLI surface; internal modules may remain for future work but are no longer documented as release commands.
- Remove unreleased provider/parity/typed-pack tools from the public MCP surface and prune the corresponding top-level Python SDK exports.
- Prune
docpull.context_packs.__all__to the public typed lanes while keeping legacy builders importable from concrete private modules for tests. - Promote
validate_pack_contract,run_context_ci,ContextCIError, andCIThresholdsthrough the root Python SDK contract. - Remove provider and observability extras from the public package extras; internal experiments that need those SDKs should install them directly.
- Remove the unreleased browser renderer runtime and keep
agent-browseras the sole local browser-rendering contract for local, Vercel, and E2B paths. - Make the
sec-filingprofile use the extractor ensemble so base installs fall back to the built-in extractor when optional Trafilatura is unavailable. - Let project mode store and sync typed source specs for packages, papers, standards, datasets, transcripts, repos, OpenAPI specs, feeds, and wiki pages.
- Bound ad hoc
docpull watchprojects to one page/depth by default and add explicit--max-pages/--max-depthwatch controls. - Let
cursor-rulesexports accept an output directory and write<skill-name>.mdcplus references, matching the CLI's file-or-directory contract. - Enrich
repo-packGitHub archive fallback with public HTML metadata when the REST API is unavailable, preserving description/topics where possible.
Fixed
- Route common raw documentation formats by extension or media type, preserve complete RFC Editor documents, and report encrypted or image-only PDFs as explicit parser/OCR requirements instead of producing misleading content.
- Honor an explicit
docpull render --live-smoke -o DIRoutput directory by preserving rendered HTML andrendered_pages.ndjsonartifacts there, while keeping bare--live-smokeruns temporary. - Return a nonzero CLI status when a crawl writes no readable records, so robots-blocked or otherwise empty public-site crawls do not look successful.
- Improve article cleanup for separated bylines, source/newsroom lines, relative timestamps, video placeholders, captions, and non-heading related story sections.
- Let
feed-packaccept bounded feeds mislabeled asapplication/octet-streamonly after normal attachment/body guards and feed shape validation. - Fix typed project source inference for repo refs containing
s, and keep HTTPS file URLs as normal web sources unless a typed lane is explicit.
[5.5.1] - 2026-06-29
Fixed
- Restore
screenshot-packcompatibility with the currentagent-browsercommand contract by falling back to batch JSON/file-output capture. - Write durable
search-packreplay config metadata for local, dry-run, budget-blocked, and provider-backed search paths.
[5.5.0] - 2026-06-29
Added
- Add typed local context packs for brand profiles, styleguide/design tokens, product/pricing extraction, schema-shaped extraction, image manifests, explicit screenshot capture, and unified local/provider search packs.
- Expose context packs across CLI, Python SDK, and MCP with durable artifacts:
result JSON, Markdown reports, source policies, citations or basis records,
replay config, pack metadata, and
run.accounting.json. - Add docs comparing DocPull's local-first context-pack boundary with hosted web intelligence APIs.
Changed
- Keep async HTTP extraction as the default path for context packs and require explicit trusted-target opt-in before screenshot/browser rendering.
- Route context-pack CSS and asset downloads through DocPull's validated HTTP transport with domain policy checks, DNS pinning, content-type checks, byte limits, and non-secret accounting.
- Update MCP plugin metadata to include the new context-pack tools.
[5.2.0] - 2026-06-24
Added
- Add environment-reference source auth for project mode, resolving credentials
into existing
AuthConfigonly at sync time and masking auth details in status, manifests, reviews, releases, and hosted payloads. - Add
docpull history,docpull review, anddocpull release context-packfor local context-repo lifecycle review and versioned release artifacts under.docpull/releases/<tag>/. - Add
docpull remote loginplus hosted remote commands for sync, status, diff, export, and release calls against a DocPull/v1API. - Add
docpull.hosted, a dependency-free Python ASGI control-plane MVP with Bearer API keys, org/project isolation, project/source CRUD, sync jobs, runs, latest diffs, context-pack export hooks, releases, webhooks, audit events, and a Postgres-ready schema contract.
Changed
- Extend the project SQLite index to track auth readiness, review summaries, and context-pack releases.
[5.1.0] - 2026-06-24
Added
- Add persistent project mode with
docpull init,add,sync,diff,status,export context-pack,eval-set, andwatchwhile preserving the legacydocpull URL ...anddocpull export PACK --format ...flows. - Add durable project run artifacts under
.docpull/runs/<run_id>/, includingrun.json,documents.jsonl,chunks.jsonl,manifest.json,errors.jsonl,accounting.json,source-health.json,documents.ndjson,corpus.manifest.json,sources.md, andlocal.pack.json. - Add the project SQLite index for sources, runs, documents, chunks, errors,
diffs, exports, and source health, with idempotent
PRAGMA user_versionschema setup. - Add deterministic project diffs with added, removed, changed, unchanged, title/path change, likely API behavior change, pricing change, and source health signals. Optional BYOK semantic summaries skip cleanly when no model key is configured.
- Add agent context-pack exports for Cursor, Claude, Codex, OpenAI, LlamaIndex, and LangChain, plus stable JSONL eval-set generation from changed or latest documents.
- Add project discovery persistence with
docpull add URL --discoveranddocpull sync --update-discovery, including exact discovered URL refreshes. - Add a project diff demo launch asset at
docs/launch-assets/docpull-project-diff-demo.png.
Changed
- Reframe README and website copy around refreshable, cited, agent-ready context packs for changing public docs.
- Clean project titles and dedupe alias pages before indexing/exporting run artifacts, reducing noisy duplicate source entries in real docs dogfood runs.
[5.0.2] - 2026-06-24
Added
- Add
docpull sharefor serving generated Markdown, HTML, or plain text reports at a simple local URL, with loopback-only binding by default and an explicit--allow-network-bindopt-in for exposed report links.
[5.0.1] - 2026-06-23
Added
- Expose non-secret run accounting receipts in MCP structured responses for
budget-blocked cloud/provider routes and local pack answers, including
run.accounting.jsonlinks when durable artifacts exist.
[5.0.0] - 2026-06-22
Added
- Add the local-first expansion surface: optional
agent-browserrendering, provider-neutral discovery packs, source policy validation, pack refresh reports, pack audits, cited local answers, JSONL/agent exports, a localhost-only pack server, authenticated-source checks, and cron-friendly local monitors. - Add a shared free-first budget contract with CLI
--budget, SDKBudgetConfig, policy-filebudget.maximum_paid_cost_usd, route explanations, stricter effective paid caps, and deterministicrun.accounting.jsonartifacts for budgeted or paid-capable runs. - Enforce fail-closed zero-dollar runs: local cache, direct HTTP, sitemap/static
discovery, local extraction, local indexing, local pack intelligence, local
monitors, and local
agent-browserrendering remain available, while live Tavily, Exa, Parallel, Vercel Sandbox, E2B, provider probes, and paid-capable benchmark routes are blocked before execution. - Add the Phase 2 zero-dollar benchmark mode and target set, including
completion classes for
complete_for_0,complete_with_local_browser,partial_for_0,requires_provider,requires_cloud_browser, andblocked_by_policy. - Add provider-free discovery scanning with
docpull discover scan URLforllms.txt, RSS/Atom feeds, OpenAPI references, richer sitemap discovery, and public GitHub documentation trees, all writing the standardcandidate_sources.ndjsondiscovery-pack contract. - Add Phase 4 escalation suggestions when local capture is partial, with local discovery/render commands first, BYOK provider dry-run/live commands next, cloud rendering last, and estimated paid request/cost guards before escalation.
- Add provider-neutral local parity workflows:
docpull extract-pack,docpull map,docpull crawl-pack,docpull research-pack, anddocpull entities-pack, plus SDK helpers and MCP tools over the same modules. These workflows write local lifecycle artifacts includingevents.ndjson,status.json,poll.report.json, andwebhook.sample.json. - Add local structured-output validation for
docpull research-pack --schemausing a dependency-free JSON Schema subset over cited local answer fields. - Add monitor lifecycle controls:
docpull monitor trigger,pause,unpause, andscheduler-snippet, plus monitor dedupe labels. - Add MCP tools for local expansion workflows:
render_url,discover_sources,fetch_discovered_sources,extract_pack,map_sources,crawl_pack,research_pack,entities_pack,refresh_pack,audit_pack,answer_pack,validate_policy,export_pack, andserve_pack_status. - Add local downstream export formats for Sheets CSV/TSV, n8n workflow JSON,
Vercel AI SDK JSON, CrewAI JSON, warehouse NDJSON, and optional Parquet via
docpull[parquet]. - Add explicit optional-renderer diagnostics:
docpull render --check,docpull render --agent-browser-bin, SDKcheck_agent_browser_availability(), doctor reporting for the externalagent-browserruntime, and a live smoke test that skips when the executable is absent. - Harden optional browser rendering by requiring HTTPS except localhost/loopback HTTP and rejecting non-default renderer action permissions.
- Add optional cloud sandbox render runtimes: Vercel through the Vercel Sandbox CLI and E2B through the E2B Python SDK/API key.
- Standardize cloud rendering on the same
agent-browser --jsoncontract as local rendering. Adddocpull render --runtime local|vercel|e2b,docpull render init ...,docpull render doctor, estimated per-render budget caps, E2B template support, prebuilt sandbox install skipping, E2B file result transport, and opt-in live cloud smoke tests gated byDOCPULL_LIVE_CLOUD_RENDER=1. - Add local pack intelligence commands:
docpull pack citations,docpull pack entities,docpull pack search, anddocpull pack brief, plus matching MCP tools for agent access to citation maps, structured signals, cited pack search, and cited briefs. - Add
docpull pack prepare,prepare_pack, and MCPpack_prepareto write the standard local pack intelligence bundle in one step. - Post-process successful Parallel, Tavily, and Exa provider context packs with local score, source-score, citation, entity, search, and brief artifacts.
- Add first-class Tavily and Exa provider adapters plus
docpull tavily ...anddocpull exa ...aliases for context-pack and extract-pack workflows. - Add a provider capability matrix and
docpull tavily map-pack, which uses Tavily Map to write a standard DocPull discovery pack. - Harden provider API key handling by rejecting unsafe key values before header use or local secret-file writes.
- Move provider capability metadata out of adapter code, split provider tests by key/adapter/CLI responsibility, and add auth path redaction for CI/agent logs.
- Add explicit provider live probes: Tavily account usage, Exa public team info, Parallel opt-in auth-gate validation, and guarded smoke probes separate from offline auth readiness.
- Add
make test-inventoryandmake test-all-localso contributors can report default pytest, fully gated pytest collection, Bun MCP, and coverage gates separately.
Changed
- Reframe public docs around DocPull as a local-first, free-first evidence engine: local and open-data routes first, BYOK providers as explicit escalation, hosted execution as a future product boundary, and no hidden paid calls, CAPTCHA bypass, stealth scraping, or proprietary web-scale index claims.
Fixed
- Resolve discovery links against the final redirected URL so moved docs sites do not produce broken relative links.
- Prefer real streamed/hidden application content over loading skeletons when extracting modern docs pages, and remove loading-only placeholders from the selected content tree.
- Clean up stdio MCP child processes in tests so the full local gate does not leak servers after client failures.
[4.4.1] - 2026-06-17
Changed
- Remove the embedded README demo GIF from the PyPI long description so the project page stays focused on install and usage content.
- Raise optional MCP dependency floors for
starletteandcryptographyso release audits install patched versions.
[4.4.0] - 2026-06-16
Added
- Add release-ready SEC filing evidence packs with
evidence.pack.json,AGENT_CONTEXT.md, validated rule confidence values, and a checked-in vendor-dependency rules example.
Changed
- Wire
docpull[proxy]to actual SOCKS proxy handling viaaiohttp-socks, keep HTTP/HTTPS proxying on aiohttp's native request path, raise the optional proxy floor toaiohttp-socks>=0.11.0, and document security-floor dependencies as Renovate-managed constraints. - Run Renovate version checks across Python, GitHub Actions, web npm, and MCP Bun manifests so dependency floor updates are raised promptly.
- Pin release build backends (
setuptools,wheel) alongsidepip,build, andtwine, and run the publish build without isolation so releases do not silently download latest build tooling.
[4.3.1] - 2026-06-15
Changed
- Tighten PyPI, GitHub, README, and website metadata around the public-web to agent-ready Markdown positioning.
- Add launch copy, comparison guidance, and marketing visibility research for developer, Python, MCP, and RAG discovery channels.
[4.3.0] - 2026-06-14
Added
- Add first-class Open Knowledge Format output via
--format okfand--profile okf, including OKF concept frontmatter, generated directoryindex.mdfiles with rootokf_version: "0.1", and docpull corpus manifests. - Add the
docpull.scraperAPI surface (Scraper,scrape_one,scrape_site, andScrapeResult) as thin scraper-native names over the existing browser-free Fetcher pipeline. - Add
docs/scraping-boundary.mdto define docpull as a local, auditable static/server-rendered web-to-context scraper rather than a general browser automation framework. - Add SQLite FTS5 indexing for
--format sqliteoutput plussearch_sqlite_documents()for local full-text retrieval. - Add static Docusaurus, Sphinx, MkDocs/Material, VitePress, Starlight, GitBook, ReadMe.io, and Redoc/Scalar-style extraction fixtures so common docs frameworks are extracted and tagged without JavaScript rendering.
[4.2.0] - 2026-06-08
Added
- Add
docpull benchmark quickfor repeatable real-site benchmark reports that compare core docpull crawls, cached reruns, and optional live Parallel Search / Search + Extract context-pack cases behind a local cost guard. - Add
docpull benchmark articleto turn benchmark JSON reports into a publishable Markdown draft with methodology, results, reproduce commands, and artifact links. - Add optional
docpull[observability]Raindrop support so benchmark cases can be emitted as metadata-only traces whenRAINDROP_WRITE_KEYis configured. - Add Raindrop event ids and per-case signals for benchmark failures, low scores, high-score cells, high-cost cells, and score-dimension warnings.
- Add Tavily and Exa live benchmark cases that normalize provider results into the same scored context-pack artifacts as core docpull and Parallel.
- Add
docpull benchmark quick --target-set tool-docs/provider-matrixfor provider-by-target matrix evals across Parallel, Exa, Tavily, Raindrop, DocPull, and low-cap adversarial public targets. The oldv2target-set name remains a compatibility alias. - Add weighted benchmark sub-scores for coverage, cleanliness, source fidelity, freshness, and density so clean and noisy targets no longer collapse to the same headline score.
- Add Tavily credit-to-dollar normalization with
--tavily-credit-usdorTAVILY_CREDIT_USD, plus a weekly GitHub Actions provider-matrix benchmark. - Add
docpull providersfor equal optional Parallel, Tavily, and Exa key status, durable key setup, and provider context-pack runs that can use any configured subset. - Add Make targets for quick, Parallel, and Raindrop benchmark runs.
Changed
- Let
docpull benchmark quick --provider auto/allrun all locally configured providers and skip missing API keys or optional SDKs without failing the core benchmark. - Write
sources.mdfor LLM-profile NDJSON output so core docpull packs score consistently with Parallel-generated context packs. - Score Parallel Search packs from their search metadata and keep fallback-pack core extraction artifacts scoped to their intended output directory.
[4.1.0] - 2026-06-07
Added
- Add optional
docpull[parallel]support for building Parallel Search + Extract context packs with local NDJSON, source Markdown, manifests, and workflow metadata. - Add
docpull parallel importfor offline fixture/demo workflows and a checked-in Parallel context-pack example fixture. - Add
docpull parallel demo, backed by a packaged fixture, so the offline context-pack demo works from an installed wheel. - Add a Parallel product cross-reference covering Search, Extract, Task, FindAll, Entity Search, Monitor, MCP, and planned follow-up workflows.
- Align Search mode choices with Parallel API docs (
turbo,basic, andadvanced) and request a Task text output schema for--task-brief. - Add source-policy, client-model, dry-run, and local cost-guard controls for live Parallel context packs.
- Add broader Parallel artifact workflows for Entity Search, FindAll, TaskGroup
batches, Monitor metadata/events, and
llms.txt/OpenAPI API packs. - Add
docpull parallel search-pack,extract-pack,task-pack,task-result, andtask-eventsfor Search-only, known-URL Extract, and Task lifecycle packs. - Add FindAll ingest, result, schema, enrich, extend, cancel, and events pack workflows.
- Add snapshot monitor creation, monitor source-policy/location/webhook/metadata controls, event-group summaries, and checked-in Parallel API-pack recipes.
- Add MCP tools for
parallel_context_pack,parallel_api_pack,pack_score, andpack_diff, plus a built-inparallelsource alias. - Add
docpull pack scoreanddocpull pack difffor local pack quality checks and refreshed-pack comparisons. - Add
docpull parallel authto check optional SDK andPARALLEL_API_KEYreadiness without storing or printing secrets. - Add raw Markdown/plain-text conversion for docs indexes such as
llms.txtthrough the normal fetch pipeline. - Add Parallel fetch policy, excerpt-size, and Search location controls to context-pack CLI and recipe workflows.
- Add Monitor list, retrieve, update, cancel, trigger, and cursor/event-group events pack workflows.
Changed
- Cap
docpull parallel context-pack --extract-limitand context-pack recipes at 20 URLs so a single Parallel Extract request stays within the documented API limit. - Make
docpull parallel taskgroup-pack --waitpoll TaskGroup status until the group is inactive before snapshotting run outputs. - Treat no-content, invalid-content, HTTP-error, and save-empty skips as failures
for
docpull --single, so single-page agent fetches do not report empty output as success. - Extend
docpull parallel runbeyond context-pack recipes so YAML/JSON recipes can dispatch the same Parallel pack workflows as the explicit CLI commands. - Refresh the documented 10,000-page benchmark wall time to the latest local audit run.
Security
- Route remote
docpull parallel api-packsources through docpull's hardened HTTPS-only URL validation, robots.txt check, DNS-pinned HTTP client, redirect revalidation, and response-size cap instead of a rawurllibfetch.
[4.0.1] - 2026-06-06
A release-readiness patch that tightens the public product boundary. No runtime API changes and no migration needed.
Changed
- Make the Python
docpull mcpserver the only documented supported MCP path for agents, plugins, Claude Code, Cursor, and Claude Desktop. - Mark the root TypeScript/Bun
mcp/tree as an internal lab, make its package metadata private, and remove end-user install instructions for that path. - Replace stale YAML example files with current CLI recipes so docs no longer
advertise removed options such as
--sources-file, TOON output,keep_variant,language, orcreate_index. - Update website examples and performance copy to match the current CLI and benchmark results.
[4.0.0] - 2026-06-04
A security + cleanup release. A multi-agent security audit closed a high-severity SSRF and nine further findings (see Security); it ships alongside a tech-debt cleanup that removes several unused public APIs (see Removed — the breaking changes that make this a major release).
Security
- DNS-rebinding TOCTOU in the URL validator (high).
UrlValidator.resolve_allowed_addresses()resolved the hostname a second time and used that unscreened answer as the connect target, so a TTL-0 attacker could pass validation with a public IP and have the socket dialed at an internal one (e.g. cloud metadata). It now resolves once and returns exactly the addresses it screened. - Wider SSRF coverage. Block CGNAT shared address space (
100.64.0.0/10) and IPv4-mapped IPv6 forms, and strip the trailing DNS root dot (localhost.) before the localhost/suffix checks — in both the Python validator and the TypeScript MCP source gate. The MCP gate additionally denies wildcard DNS-rebinding hosts (*.nip.io,*.sslip.io,*.xip.io). - robots.txt memory-exhaustion DoS. Cap the robots.txt body read at 512 KB, matching the existing sitemap limit.
- YAML frontmatter injection. Frontmatter list items (tags/keywords sourced from page JSON-LD / OpenGraph) are quoted, escaped, and stripped of CR/LF so a hostile page cannot inject top-level frontmatter keys.
- Conditional-request header injection. Cached
ETag/Last-Modifiedvalues are stripped of CR/LF/NUL before being reused asIf-None-Match/If-Modified-Since. - Supply chain. Pin release tooling (
pip/build/twine) viarequirements-release.txt; drop six unused (ghost) dependencies from the MCP package; bumpaiohttpto>=3.14.0(CVE-2026-34993, CVE-2026-47265).
Removed
- Unused public methods on
CacheManager(breaking for any external caller):has_changed,is_fetched,is_failed,get_failed_urls,get_cache_stats,clear_state, andhas_resume_datahad no callers in the library or tests. Incremental fetch and resume are unaffected — they use the retainedupdate_cache,mark_fetched,mark_failed,get_fetched_urls,get_pending_urls,save_/load_/clear_discovered_urls, andevict_expired. StreamingDeduplicator.is_duplicate(breaking): unused read-only probe. Usecheck_and_register, whose first return value reports whether content was new.DocpullConfig.from_yaml_file(breaking): unused convenience wrapper. UseDocpullConfig.from_yaml(path.read_text()).
Changed
- Internal cleanup with no API or behaviour change: removed dead code (the unused
concurrencypackage,logging_config, and several private dead methods) and de-duplicated the discovery HTML-fetch helper and the HTTP GET/HEAD redirect re-validation path.
Fixed
- MCP indexing of large libraries. pgvector embedding inserts are now batched under PostgreSQL's 32767 bind-parameter ceiling, so a library with thousands of chunks indexes in one transaction instead of failing.
[3.0.2] - 2026-05-29
A small release hygiene patch for the MCP hardening release. No API changes; no migration needed.
Fixed
- Runtime version now matches package metadata.
docpull --versionand the default HTTPUser-Agentreport3.0.2instead of the stale3.0.0value that remained indocpull.__version__after the 3.0.1 publish.
Tests
- Added a stdio MCP smoke test. The test starts
docpull mcpthrough the official MCP client, verifies the advertised 8-tool surface, checks structuredlist_sourcesoutput, and confirms SSRF rejection still flows through a real MCPcall_toolrequest.
[3.0.1] - 2026-05-29
A security and correctness patch. No API changes; no migration needed.
Security
- User-defined MCP sources are validated on load. Entries in
~/.config/docpull-mcp/sources.yamlare now rejected unless the name is a safe identifier, the URL is HTTPS to a public host (private, loopback, link-local, and internal-suffix hosts are blocked), andmax_pagesis in range. Previously a hand-edited config could pointensure_docsat an internal address. grep_docsbounds regex execution per line. On top of the existing total wall-clock budget, each line now matches under a per-line timeout, closing the remaining catastrophic-backtracking (ReDoS) window for a pathological pattern against a single long line.
Fixed
- Cache timestamps are timezone-aware UTC. Persisted timestamps (cache manifest, save steps, MCP metadata) use UTC consistently; legacy naive timestamps are parsed as UTC so cache-TTL comparisons stay deterministic instead of mis-expiring entries.
- Swallowed exceptions are now logged. robots.txt parsing, the OpenAPI and SPA heuristics, and link extraction log skipped or invalid input at debug level instead of silently dropping it.
[3.0.0] - 2026-04-26
The deprecations 2.4 promised. Six config fields that have emitted a
DeprecationWarning since 2.4 are now gone, and the naming_strategy
literal no longer accepts the "flat" / "short" aliases that were
documented as "aliased to 'full' until 3.0".
Breaking
ContentFilterConfigremoved fields —language,exclude_languages,deduplicate,max_total_size,exclude_sections. All have been no-ops since 2.4 with a deprecation warning on use; pydantic will now reject configs that set them (model_config = {"extra": "forbid"}). Fordeduplicate=True, switch tostreaming_dedup=True. The other fields had no replacement because they had no effect.OutputConfig.create_indexremoved — also a no-op since 2.4. Drop the field from your config; nothing to migrate.OutputConfig.naming_strategyliteral narrowed — the alias values"flat"and"short"(which silently behaved like"full") are no longer accepted. Use"full"directly."hierarchical"is unchanged.docpull.deprecatedlogger removed — the dedicated logger and the per-callDeprecationWarninginfrastructure for the above fields are gone with them. Filters that targeteddocpull.deprecatedcan be removed.
Migration
If your config file or DocpullConfig(...) call sets any of the
removed fields, delete those lines. Pydantic's forbid policy will
otherwise raise ValidationError at construction time with a clear
"Extra inputs are not permitted" message naming the field.
[2.5.1] - 2026-04-25
A small but real bugfix: the grep_docs → read_doc round-trip was
broken. grep_docs returned paths with the library name prepended
(e.g. hono/middleware/basic-auth.md), but read_doc joins
library and path itself, so passing a grep result verbatim
produced hono/hono/middleware/basic-auth.md and 404'd. The
contract advertised in the read_doc description ("the natural
follow-up to grep_docs: pass the library + path it returned")
didn't actually work.
Fixed
grep_docsreturns library-relative paths. Each result in the structuredfilespayload now has bothlibrary(the library name) andpath(relative to the library root). Pass them straight intoread_doc(library=..., path=...)— no munging. Human-readable text rendering still showslibrary/pathas the qualified identifier, so existing terminal output looks identical.- Tool descriptions for
grep_docsandread_docupdated to match the actual contract.
Schema
_GREP_DOCS_OUTPUT_SCHEMA.files.itemsnow requireslibraryin addition topath. Existing consumers that readpathwill get a different (now correct) value; consumers that don't pipe grep results intoread_docare unaffected.
Tests
- Added
test_grep_to_read_doc_roundtripandtest_grep_to_read_doc_roundtrip_with_line_sliceregression tests that passlibraryandpathfrom grep verbatim intoread_docand assert success. - Added
test_grep_docs_path_is_library_relative_in_subdirto cover nested files. - Updated
test_grep_docs_structured_payloadto assert the newlibraryfield and exact (not just suffix-matched) path value.
[2.5.0] - 2026-04-25
A focused MCP-server hardening pass. Closed three exploitable security
holes in the agent-facing tools, added the missing ToolAnnotations
that gate Anthropic Directory submission, exposed structured output
alongside the rendered text on every tool that carries data, and
added the three tools an agent obviously wants — read_doc to follow
up a grep_docs hit, plus add_source / remove_source to manage
the user registry programmatically.
Added
read_doc(library, path, line_start?, line_end?)— read a Markdown file from a fetched library, optionally line-sliced. The natural follow-up aftergrep_docsreturns a hit; agents no longer need filesystem access for surrounding context. Path is resolved and confirmed to stay under the library root.add_source(name, url, ...)— add or update a user source alias in the writablesources.yaml. Refuses to shadow a builtin alias unlessforce=true; URL is HTTPS-only and validated against the same SSRF rules asfetch_url. Atomic write (tmp + rename).remove_source(name, delete_cache?)— remove a user source alias and optionally its cached Markdown directory. Cannot remove builtins (suggestadd_source(force=true)to shadow instead). Cache deletion does a defense-in-depth resolved-path check.ToolAnnotationson every tool —readOnlyHint/destructiveHint/idempotentHint/openWorldHint/title. Required for Anthropic Directory submission and unlocks host auto-approve for the four read-only tools.- Server
instructions— system-prompt hint telling agents the call ordering (list_sources → ensure_docs → grep_docs → read_doc). - Progress notifications —
ensure_docsforwardsFETCH_COMPLETEDevents as MCP progress to clients that supplied aprogressTokenon the call. - Structured output (
outputSchema+structuredContent) onlist_sources,list_indexed,grep_docs,read_doc,ensure_docs,add_source,remove_source. Clients that consumestructuredContentget parseable JSON; clients that don't still see the rendered Markdown text.
Fixed
- SSRF in
fetch_url— schema previously accepted any string with no scheme/host enforcement. An agent could requesthttp://169.254.169.254/,http://localhost,file:///etc/passwd, etc. Now validated upfront with the sameUrlValidator(HTTPS-only, no localhost / private / link-local IPs) the crawler uses, instead of relying on the slow pipeline error path. - Path traversal in
grep_docs/read_docvialibrary—docs_dir / librarydid not validatelibrary, solibrary="../../etc"walked anywhere the process could read.read_doc'spatharg was similarly unchecked. Both now reject unsafe names (is_safe_library_name) andread_docadditionally resolves the joined path and confirms it stays under the library root. - ReDoS in
grep_docs— pattern was compiled with no length cap and run line-by-line over every cached.md; Pythonrehas no timeout knob. Now cap pattern length at 1000 chars and apply a 10s wall-clock budget across files. isErrorflag was being silently dropped — the previous_call_toolreturned a barelist[TextContent], which the SDK's legacy path hardcodes asisError=Falseregardless of what the handler intended. Every error your tools raised was being reported to clients as success. Now_call_toolreturnsCallToolResultdirectly soisErrorpropagates correctly.ensure_docspartial-fetch detection — a crash mid-crawl used to leave files on disk with no meta, and the next call would re-fetch (correct, but wasteful) or — if a stale meta from a prior run was present — trust the half-fetched cache. Meta writes are now atomic (tmp + rename) and apartial=trueflag marks half-fetches so_cache_freshtreats them as stale.grep_docshonorscontext > 1— the schema advertisedmaximum: 3but the implementation only ever rendered one line either side. Now renders up tocontextlines on each side.load_user_sourcessilently swallowed YAML errors — a typo in the user'ssources.yamlproduced "Unknown source" instead of surfacing the parse failure. Now logs a warning at WARNING level.
Changed
- Tighter input validation in the MCP
_call_tooldispatcher: required strings checked with_require_str, ints coerced with_coerce_int. Errors that used to surface as ugly "invalid literal for int(...)" now return clear messages naming the bad argument. - Tighter input schemas:
https://pattern onfetch_url.url,enumoncategory, regex +maxLengthonlibraryeverywhere it appears,maxLength: 1000ongrep_docs.pattern, integer bounds onmax_tokens/max_pages. _cache_freshnow also requires the source directory to contain at least one.mdfile — a manually-rm -rf'd cache no longer reports as fresh._PROFILE_ALIASESmapping deleted;_resolve_profilenow goes throughProfileNamedirectly, eliminating drift.- 39 new MCP tests (61 total in
test_mcp_tools.py, 316 in the full suite). Coverage includes SSRF rejection, path traversal, oversized regex, partial-meta freshness, structured payloads, and the new write tools.
[2.4.0] - 2026-04-26
A two-pass cleanup. The first pass closed every claim the code didn't back
(seven scaffolded-but-unwired config fields, the no-op fence-language regex,
the "ETag-based caching" claim that never sent If-None-Match, the
robots.txt UA mismatch, the MdxSourceExtractor dead branch). The second
pass earned the local-first / agent-native / zero-trust pitch the marketing
points at: cookie-banner stripping, rich frontmatter, --skill mode,
streaming discovery, conditional GET, and a measured 10k-page benchmark.
Added
- Conditional GET on cached pages:
FetchStepnow sendsIf-None-MatchandIf-Modified-Sincefrom the manifest, and a304 Not Modifiedresponse short-circuits withSkipReason.CACHE_UNCHANGED. Previously the marketing claimed ETag-based skipping but the headers were never actually sent. Re-runs against an unchanged site now transfer near-zero bytes. - Hierarchical naming:
output.naming_strategy: hierarchical(set by the Mirror profile) preserves URL paths as nested directories (/api/auth/oauth2→api/auth/oauth2.md), with sanitized segments and trailing-slash →index.mdcollapse. Path-traversal segments (..) are neutralized so URL-driven escapes can't leave the output dir. --skill NAMEgenerates agent-ready skills/rules:docpull URL --skill foo --skill-agent allwrites scraped pages under.docpull/skills/foo/references, creates Claude Code and CodexSKILL.mdwrappers, and writes a Cursor.cursor/rules/foo.mdcproject rule. With--output-dir, the corpus is staged there while explicit agent targets still write active wrappers. The manifestdescriptionis derived from the first page's OpenGraph or JSON-LD metadata, with--skill-descriptionavailable as an explicit override.--require-pinned-dnsrefuses proxy configurations that delegate DNS to the proxy. Before this, running with--proxysilently weakened the SSRF posture (only a startup warning was logged); the new flag makes the trade-off explicit. Default off; intended for agent-driven workflows.--no-streaming-discovery: backstop flag for the new producer- consumer fetch pipeline (see Changed). Falls back to the legacy discover-all-then-fetch behavior in case backpressure regressions surface in the wild.max_file_sizecontent-filter is now wired to the HTTP client's per-response cap. Previously hardcoded at 50 MiB; users can now lower it for OOM-prevention on runaway responses.- Rich frontmatter: every Markdown file now ships with a heading
outline (top-level
h1/h2, ≤12 entries), an ISO 8601crawled_attimestamp, OpenGraphdescription, and a whitelisted slice of JSON-LD/microdata fields (author,published_time,keywords, etc.). Previously OG/JSON-LD extraction ran but the result was dropped. - MCP surface polish:
ensure_docsaccepts aprofileargument (rag/mirror/quick/llm);grep_docsranks results by per-file match density and renders ±1 line of context per hit (configurable viacontext);list_indexedreports humanized fetch age per source;fetch_urlincludes chunk count in its response header. - 10,000-page benchmark:
tests/benchmarks/test_10k_pages.pystands up a synthetic localhost site with injected duplicates and reports wall time, peak RSS delta, manifest size, p50/p95/p99 per-page latency, and time-to-first-save. Gated behindDOCPULL_BENCHMARK_10K=1. README's new## Performancesection documents the headline numbers.
Fixed
- Code-fence language normalization: html2text emits
[code]…[/code]blocks without language tags by default, and the post-conversion regex meant to fix this was a self-replace no-op. Pages with Prism (class="language-python"), highlight.js (lang-py/hljs-language-X), Shiki, or GitHub-style (highlight-source-rust) syntax classes now produce GFM fenced blocks with the right language tag.plaintext/text/noneare correctly treated as "no language." - Cookie / consent banner leakage: the FAQ claimed common banners
were stripped, but no selectors targeted the vendor SDK shapes. Added
selectors for OneTrust, Osano, Cookiebot, CookieLaw, CookieConsent,
Iubenda, Termly, and generic
.cookie-*/.gdpr-*/.consent-*patterns plusaria-label*="cookie|consent|gdpr"fallbacks. Pages that legitimately discuss cookies in their body are unaffected (selectors are structural, not text-based). - Streaming dedup hashed full Markdown including frontmatter:
meant two URLs serving byte-identical body content never deduped
because their
source:and (after this release)crawled_at:fields differed. Dedup now strips frontmatter before hashing — the point of streaming dedup is "same body content," not "same bytes." - Robots.txt UA mismatch: docpull matched robots.txt rules as
docpull/2.0while sending requests asMozilla/5.0 ... AppleWebKit .... Site operators scoping rules atUser-Agent: docpullgot no effect. Both surfaces now use the same UA, derived from the HTTP client. - Empty
cache-freshshort-circuit when output is missing: if a user clearedoutput/but keptcache/, the new conditional GET would have skipped on 304 and left no Markdown on disk.FetchStepnow suppresses conditional headers when the expected output file is absent, forcing a fresh re-fetch.
Changed
- Default User-Agent:
docpull/{version} (+https://github.com/raintree-technology/docpull), replacing the previous Mozilla camouflage. Sites that whitelisted Mozilla patterns may need to allow the new UA;--user-agentcontinues to override. The pitch is "polite crawler" — disguising as a browser contradicted that. - Streaming discovery → fetch (default): URLs now flow through a
bounded worker pool as the discoverer yields them, instead of being
collected to a list before any fetching begins. First
PAGE_SAVEDon a 10,000-page synthetic site now fires within ~70 ms of the run starting (vs. waiting for full discovery before). The discoverer awaits whenurl_queueis full, so backpressure self-regulates.--no-streaming-discoveryfalls back to the legacy path. - Mirror profile keeps flat naming by default in 2.x to preserve
existing users' output paths. Hierarchical is now opt-in via
--naming-strategy hierarchical(CLI) oroutput.naming_strategy(YAML). The Mirror profile default flips to hierarchical in 3.0; the upgrade will be flagged in the 3.0 release notes. MdxSourceExtractorremoved fromDEFAULT_CHAIN(it always returnedNone). The class andfind_mdx_source_urlhelper are still exported for callers that want to wireprefer_sourcemanually.
Deprecated
The following config fields warn at runtime when set to a non-default
value and will be removed in 3.0. Each was scaffolded but never read by
the pipeline; CLI flags backed by them have been dropped from --help.
content_filter.languageand--language(no language detector ever shipped — pursue if a real user asks)content_filter.exclude_languagescontent_filter.deduplicate(post-processing dedup;streaming_dedupcovers the use case)content_filter.exclude_sectionscontent_filter.max_total_size(cumulative byte budget across an async fetch is racy; per-pagemax_file_sizeis the right knob)output.create_index(noINDEX.mdgenerator; downstream tools don't need one)
Security
- Honest UA disclosure: see Changed. Polite-crawling claim now lines up with what site operators see in their logs.
--require-pinned-dns: see Added. Closes a previously-undisclosed gap where--proxydisabled docpull's connector-level DNS pinning.- Hierarchical-naming traversal: URL path segments are sanitized
(
..→index, special chars →_, runs of underscores collapsed) before being joined into the output path. TheSaveStepbase-dir guard remains the second line of defense.
[2.3.0] - 2026-04-24
Sharpened positioning around the agent / RAG use case, plus real bug fixes surfaced by validation against Next.js, Supabase, Anthropic, FastAPI, Tailwind, and Drizzle documentation sites.
Added
- Framework-specific fast extractors: Next.js
__NEXT_DATA__, Mintlify, OpenAPI / Swagger JSON rendered directly to Markdown, plus source-type tagging for Docusaurus and Sphinx. Runs before the generic extractor. - Next.js App Router detection via
self.__next_f.push, router state tree, and/_next/static/path markers — no longer relies on__NEXT_DATA__, which is absent on modern App Router pages. - SPA detection (pre- and post-conversion): pages that produce only
Loading...shells are skipped with a clear reason.--strict-js-requiredturns this into a hard error for agents that want to route elsewhere. - Trafilatura extractor as an optional alternative content extractor
(
pip install docpull[trafilatura], then--extractor trafilatura). - Token-aware Markdown chunking:
--max-tokens-per-file Nsplits pages on heading then paragraph boundaries. Exact counts withtiktoken, character-estimate fallback otherwise. - NDJSON output format (
--format ndjson) for streaming one record per page or per chunk.--streamwrites to stdout for live pipeline consumption. llmprofile: bundles NDJSON + 4k-token chunks + rich metadata + dedup.--single/fetch_one(url): fast single-page path with no discovery, designed for AI-agent tool loops.- Python MCP server (
docpull mcp): exposesfetch_url,ensure_docs,list_sources,list_indexed, andgrep_docstools over stdio. Install viapip install docpull[mcp].
Fixed
- robots.txt redirect handling: Cloudflare/HTTP-2 responses send
lowercase header names, but the
Locationlookup was case-sensitive, causing 301/308 redirects to be treated as errors. This blockeddocs.anthropic.comand any other site whose robots.txt was redirected. - html2text link escape artifacts: cleaned up mangled Markdown link targets
containing
prefix/<https:/real.url>in the post-processing pass; handles both text and image-only (empty-text) links.
Removed
- Dead dependencies:
requests(replaced byaiohttpin v2.0) andgitpython(never used in v2+).
Changed
ContentFilterConfiggainsextractor,enable_special_cases, andstrict_js_requiredfields.OutputConfiggainsmax_tokens_per_file,tokenizer,emit_chunks, andndjson_filename.
2.0.0 - 2025-11-29
Breaking Changes
- Complete architecture rewrite with new Python API
- Moved to
src/layout (PEP 517/518 compliant) - Old
GenericAsyncFetcherreplaced byFetcherclass with async context manager - Configuration now uses Pydantic models (
DocpullConfig)
Added
- Streaming event API: Async iterator interface for real-time progress tracking
- CacheManager: Persistent caching with O(1) lookups, batched writes, TTL eviction
- StreamingDeduplicator: Real-time duplicate detection during fetch
- Profiles: Built-in
rag,mirror,quickprofiles with sensible defaults - CLI cache options:
--cache,--cache-dir,--cache-ttl,--no-skip-unchanged - Pipeline architecture: Modular steps (Validate, Fetch, Convert, Dedup, Save)
Changed
- Cache uses sets internally for O(1) URL membership checks
- Consistent SHA-256 hashing across cache and dedup (accepts str or bytes)
- ETag and Last-Modified headers now extracted and cached
Removed
- Old fetcher classes (
GenericAsyncFetcher,AsyncDocFetcher, etc.) DedupTrackerreplaced byStreamingDeduplicator- Legacy config fields (
incremental,update_only_changed)
1.5.0 - 2025-11-28
Added
- Proxy support: HTTP, HTTPS, SOCKS5 via
--proxyorDOCPULL_PROXYenv var - Retry with exponential backoff:
--max-retries,--retry-base-delayfor transient failures - Better encoding detection: Intelligent charset detection for international docs
- URL normalization: Reduces duplicate fetches by 10-20%
- Content hash change detection: SHA-256 hashing for efficient incremental updates
- Custom User-Agent:
--user-agentflag
Changed
- robots.txt compliance is now mandatory (cannot be disabled)
- Automatically respects Crawl-delay directives
1.4.0 - 2025-11-28
Breaking Changes
- Removed profile system entirely - use URLs directly
--sourceflag removed; use positional URL arguments- Python API:
urlparameter instead ofurl_or_profile
1.3.0 - 2025-11-20
Added
- Rich metadata extraction:
--rich-metadataextracts Open Graph, JSON-LD, microdata - Enhanced frontmatter with author, description, keywords, images, publish dates
Changed
- Removed 7 built-in profiles; generic fetcher works for all sites
1.2.0 - 2025-11-16
Added
15 major features for optimization and workflow automation:
Optimization
--language/--exclude-languages: Filter by language--deduplicate/--keep-variant: Remove duplicate files--max-file-size/--max-total-size: Size limits--exclude-sections: Remove verbose sections
Output
--format: markdown, toon, json, sqlite--naming-strategy: full, short, flat, hierarchical--create-index: Generate INDEX.md
Workflow
--sources-file: Multi-source YAML configuration--incremental/--update-only-changed: Resume and update detection--git-commit/--git-message: Git integration--archive/--archive-format: Create archives--post-process-hook: Python plugin system
Changed
- PyYAML and GitPython now required dependencies
1.1.0 - 2025-11-14
Added
--doctorcommand for installation diagnostics- TROUBLESHOOTING.md documentation
1.0.0 - 2025-11-07
Added
- Initial release
- Async + parallel fetching
- Security: HTTPS-only, path traversal protection, XXE protection, size limits
- Rate limiting and timeout controls
- YAML frontmatter in output files