Changelog

September 5, 2026 · View on GitHub

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]

Fixed

  • Print only the exact pinned pipx install command when the plugin launcher cannot execute DocPull.

[6.5.5] - 2026-09-04

Fixed

  • Report the DocPull package version during MCP initialization and attach the advertised cost description and docpull/cost metadata to every MCP tool.

[6.5.4] - 2026-09-04

Fixed

  • Treat a non-executable docpull path as an unavailable runtime and return the same actionable setup guidance as a missing executable.

[6.5.3] - 2026-09-04

Fixed

  • Shorten the Codex activation prompt to the host's 128-character limit so the focused research trigger is loaded instead of ignored.

[6.5.2] - 2026-09-04

Fixed

  • Add a dependency-free plugin launcher that reports the exact pipx setup command when the docpull executable is unavailable.

[6.5.1] - 2026-09-04

Fixed

  • Constrain the optional MCP runtime to the compatible 1.x series. This prevents fresh docpull[mcp] installations from resolving MCP 2.x, whose low-level server API is incompatible with DocPull 6.5.

[6.4.0] - 2026-07-16

Added

  • Add website-pack across CLI, Python workflow API, and MCP. The workflow emits the pinned website.snapshot.v1 schema, stable URL-based document identities, content-hash versions, page roles, source authority, OKF/raw representations, bounded optional visuals, portable-v3 manifests, and recursively verifiable artifact hashes.
  • Add verified-baseline diff states for added, changed, unchanged, removed, failed, and blocked pages, plus golden/tamper fixtures.

Changed

  • Let brand, product, policy, styleguide, image, and relationship workflows consume existing packs without network fallback. Product extraction now keeps trial metadata separate from ordinary price fields and excludes testimonial copy from feature evidence.

Removed

  • Remove the standalone DocPull website, its Next.js workspace, and web-only CI/release gates. GitHub, PyPI, and the MCP Registry remain the canonical public product surfaces.

[6.3.0] - 2026-07-16

Added

  • Add relationship-pack across CLI, SDK, MCP, workflow registry, and project sources. It emits cited owned_by, operated_by, acquired_by, franchised_by, and invested_in observations plus one coverage result per input; missing evidence remains a coverage gap and never becomes an independence claim.
  • Put core fetch, crawl, and dataset-pack on WorkflowRequest/WorkflowResult; empty and partial crawls now retain a structured current-run manifest, progress, budget usage, hashes, replay configuration, warnings, and typed failures.
  • Support bounded remote HTTPS JSON/CSV dataset snapshots with original and resolved URL provenance, query parameters, content type, and deterministic SHA-256 hashes.
  • Add the DocPull product site with home, pricing, privacy, terms, llms.txt, robots, sitemap, manifest, branded icons, and reusable launch assets.
  • Add metadata-only native integration context adapters that produce cited, rights-labeled document records without storing provider secrets or raw customer payloads.
  • Add a committed uv.lock, pinned Mise runtimes, Next.js ESLint gates, and dependency review for reproducible local and CI environments.

Changed

  • Make CLI success strictly current-run scoped by default. Add explicit --exit-policy usable-output for consumers that intentionally accept older records in a shared directory.
  • Classify authority per entity/source in multi-company bundles and add official_corporate, government_registry, regulatory_filing, press_release, and local_reporting roles.
  • Extract relationship observations into intelligence bundles while keeping every candidate unresolved until downstream human review.
  • Load the root SDK and internal package exports lazily while preserving every documented import and CLI, Python SDK, and MCP surface.
  • Coalesce byte-equivalent concurrent HTTP GETs without retaining completed responses, and isolate requests with different timeouts or headers.
  • Journal frontier transitions between compact snapshots so interrupted runs remain resumable without rewriting the full frontier after every URL.
  • Reuse pack reads, citation analysis, entity extraction, and graph indexes across intelligence workflows; use constant-time document/source lookups and bounded top-result selection for local search.
  • Centralize schema versions and export-format registries in lightweight modules to reduce import cost and duplicated contract definitions.

Fixed

  • Populate retryable acquisition failures with stable codes such as http_429, fetch stage, HTTP status, attempt count, and Retry-After seconds.
  • Prevent stale records in a reused output directory from turning a zero-record current run into exit status 0.
  • Restore automatic CI, CodeQL, security, benchmark, live-smoke, and metrics triggers after the runtime-tooling migration, and make the Python matrix select the advertised patch versions explicitly.
  • Extend release-readiness checks to lint, type-check, audit, and build the web workspace, and repair stale repository links in issue templates.

Compatibility

  • Keep all 6.2 pack, workflow, CLI, SDK, MCP, schema, and tracker contracts readable and importable; 6.3 changes are additive or internal optimizations.

[6.2.0] - 2026-07-16

Added

  • Define DocPull as the evidence and acquisition engine and publish versioned WorkflowRequest, WorkflowResult, ArtifactManifest, intelligence.bundle.v1, and ChangeEvent contracts with bundled JSON Schemas.
  • Add generic progress, warnings, failures, budget usage, SHA-256 artifact hashes, pack/run identity, and scheduler-neutral replay settings.
  • Promote brand, product, styleguide, image, screenshot, and new policy packs to public CLI, SDK, and MCP workflows, plus declarative project source types for brand, product, styleguide, visual, and policy evidence.
  • Add policy discovery/classification, effective dates, stable clause hashes, and clause-level before/after candidates without legal conclusions.
  • Add deterministic tracker bundles with source snapshots, document versions, precise evidence spans, observations, confidence, evidence strength, source authority, warnings, and change candidates.
  • Add idempotent change events separating structural, textual, and semantic candidates for pricing, positioning, product, security, and policy changes.

Changed

  • Enrich product pricing extraction with plans, currencies, billing intervals, trials, feature gating, page-text provenance, and incidental/visual price exclusion.
  • Make JSON-only eval preparation retain the required PACK_CARD.md artifact.
  • Keep company_brain.bundle.json and its SDK builder as compatibility aliases for intelligence.bundle.v1.

Compatibility

  • Validate old packs remain readable and retain legacy result filenames while new workflow sidecars and schema fields are additive.

[6.1.0] - 2026-07-14

Added

  • Add local pypdf document extraction through the pdf extra, including isolated subprocess execution, resource and output limits, secure temporary files, structural PDF classification, and extraction-quality provenance.
  • Add scorer v4, portable report schema v3, explicit comparison scopes, encrypted evidence escrow, signed publication verification, and clean-wheel subject provenance to the isolated benchmark lab.

Changed

  • Prefer pypdf for remote PDF extraction before optional MarkItDown and Unstructured fallbacks, with consistent CLI, Python, and MCP controls.
  • Fence structured raw responses with adaptive labeled Markdown fences while preserving Markdown and documentation-style plain text directly.
  • Make repository-hosted manual benchmark runs exploratory only; comparative claim evidence must come from the external sealed-holdout process.

Fixed

  • Reject fused words in benchmark evidence while retaining explicit line-break hyphen repair and punctuation-separated token equivalence.
  • Recompute benchmark summaries and billing from observations, rejecting forged, incomplete, duplicate, or tampered report and publication data.

[6.0.0] - 2026-07-01

Added

  • Add DocPull output contract v3 validation with docpull pack validate, raw/agent/eval levels, v3 manifests, raw sidecars for file-backed outputs, record-level citations, chunk-safe SQLite output, and listing item sidecars for link-dense event/news pages.
  • Add docpull parse, a local document parsing lane that emits v3 packs from plain text/Markdown directly and optional MarkItDown or Unstructured parser backends for complex office/PDF files.
  • Add explicit --remote-documents pdf support for locally parsing fetched PDFs while keeping remote documents, browsers, and cloud parsing disabled by default.
  • Add optional Presidio-backed PII detection for docpull pack audit --redaction and docpull pack redact, while keeping deterministic regex redaction as the default backend.
  • Add docpull openapi-pack to convert local or HTTPS OpenAPI JSON/YAML specs into v3 endpoint and component-schema records.
  • Add docpull feed-pack to convert RSS, Atom, JSON Feed, or pages that advertise feeds into item-level v3 records with feed item sidecars, listing sidecars, and freshness metadata.
  • Add typed knowledge lanes for known-source context dependencies: docpull paper-pack, repo-pack, package-pack, standards-pack, dataset-pack, transcript-pack, and wiki-pack. These commands emit v3 raw packs, validate with docpull pack validate, prepare to agent/eval grade, and export through the existing pack export formats without adding new MCP tools. Remote typed lanes also support opt-in metadata caching, typed sidecar roots, async SDK wrappers, standards section records, exact streamable dataset row counts, and explicit official API source contracts for arXiv, Crossref, and NCBI E-utilities.
  • Add --extractor ensemble, which scores built-in and optional trafilatura extraction candidates and keeps the strongest Markdown output.
  • Add docpull ci, a local Context CI gate for project and standalone pack context. It validates lockfiles, current score/audit sidecars, coverage, citation coverage, eval-grade artifacts, rights constraints, and optional context prediction pass rates while writing context-ci.report.json and CONTEXT_CI.md.
  • Add GitHub Actions workflow definitions for scheduled free live smokes over typed official APIs and ordinary web sources, so source drift can be checked outside the default offline test suite.
  • Add scripts/real_feature_smoke.py, an opt-in real-data acceptance harness for the free/local public surface with optional cloud-render lanes.
  • Document the Context CI GitHub Actions workflow, context-pack contract, and hidden-eval human review protocol.
  • Add Context CI examples, Context Pack Contract v3, a design-partner playbook, and a spec-only hosted control-plane extension.
  • Make current-context-qa the default evalgen task type while accepting current-docs-qa as a legacy alias.
  • Add a Context CI benchmark report over local demo packs showing passing, manifest-mismatch, coverage, rights, citation, and stale-answer gates.

Changed

  • Consolidate the unreleased v3 surface around context dependencies, pack validation/preparation, exports, Context CI, document parsing, OpenAPI packs, and the canonical agent-browser renderer contract.
  • Remove unreleased top-level parity/provider/benchmark commands from the public CLI surface; internal modules may remain for future work but are no longer documented as release commands.
  • Remove unreleased provider/parity/typed-pack tools from the public MCP surface and prune the corresponding top-level Python SDK exports.
  • Prune docpull.context_packs.__all__ to the public typed lanes while keeping legacy builders importable from concrete private modules for tests.
  • Promote validate_pack_contract, run_context_ci, ContextCIError, and CIThresholds through the root Python SDK contract.
  • Remove provider and observability extras from the public package extras; internal experiments that need those SDKs should install them directly.
  • Remove the unreleased browser renderer runtime and keep agent-browser as the sole local browser-rendering contract for local, Vercel, and E2B paths.
  • Make the sec-filing profile use the extractor ensemble so base installs fall back to the built-in extractor when optional Trafilatura is unavailable.
  • Let project mode store and sync typed source specs for packages, papers, standards, datasets, transcripts, repos, OpenAPI specs, feeds, and wiki pages.
  • Bound ad hoc docpull watch projects to one page/depth by default and add explicit --max-pages / --max-depth watch controls.
  • Let cursor-rules exports accept an output directory and write <skill-name>.mdc plus references, matching the CLI's file-or-directory contract.
  • Enrich repo-pack GitHub archive fallback with public HTML metadata when the REST API is unavailable, preserving description/topics where possible.

Fixed

  • Route common raw documentation formats by extension or media type, preserve complete RFC Editor documents, and report encrypted or image-only PDFs as explicit parser/OCR requirements instead of producing misleading content.
  • Honor an explicit docpull render --live-smoke -o DIR output directory by preserving rendered HTML and rendered_pages.ndjson artifacts there, while keeping bare --live-smoke runs temporary.
  • Return a nonzero CLI status when a crawl writes no readable records, so robots-blocked or otherwise empty public-site crawls do not look successful.
  • Improve article cleanup for separated bylines, source/newsroom lines, relative timestamps, video placeholders, captions, and non-heading related story sections.
  • Let feed-pack accept bounded feeds mislabeled as application/octet-stream only after normal attachment/body guards and feed shape validation.
  • Fix typed project source inference for repo refs containing s, and keep HTTPS file URLs as normal web sources unless a typed lane is explicit.

[5.5.1] - 2026-06-29

Fixed

  • Restore screenshot-pack compatibility with the current agent-browser command contract by falling back to batch JSON/file-output capture.
  • Write durable search-pack replay config metadata for local, dry-run, budget-blocked, and provider-backed search paths.

[5.5.0] - 2026-06-29

Added

  • Add typed local context packs for brand profiles, styleguide/design tokens, product/pricing extraction, schema-shaped extraction, image manifests, explicit screenshot capture, and unified local/provider search packs.
  • Expose context packs across CLI, Python SDK, and MCP with durable artifacts: result JSON, Markdown reports, source policies, citations or basis records, replay config, pack metadata, and run.accounting.json.
  • Add docs comparing DocPull's local-first context-pack boundary with hosted web intelligence APIs.

Changed

  • Keep async HTTP extraction as the default path for context packs and require explicit trusted-target opt-in before screenshot/browser rendering.
  • Route context-pack CSS and asset downloads through DocPull's validated HTTP transport with domain policy checks, DNS pinning, content-type checks, byte limits, and non-secret accounting.
  • Update MCP plugin metadata to include the new context-pack tools.

[5.2.0] - 2026-06-24

Added

  • Add environment-reference source auth for project mode, resolving credentials into existing AuthConfig only at sync time and masking auth details in status, manifests, reviews, releases, and hosted payloads.
  • Add docpull history, docpull review, and docpull release context-pack for local context-repo lifecycle review and versioned release artifacts under .docpull/releases/<tag>/.
  • Add docpull remote login plus hosted remote commands for sync, status, diff, export, and release calls against a DocPull /v1 API.
  • Add docpull.hosted, a dependency-free Python ASGI control-plane MVP with Bearer API keys, org/project isolation, project/source CRUD, sync jobs, runs, latest diffs, context-pack export hooks, releases, webhooks, audit events, and a Postgres-ready schema contract.

Changed

  • Extend the project SQLite index to track auth readiness, review summaries, and context-pack releases.

[5.1.0] - 2026-06-24

Added

  • Add persistent project mode with docpull init, add, sync, diff, status, export context-pack, eval-set, and watch while preserving the legacy docpull URL ... and docpull export PACK --format ... flows.
  • Add durable project run artifacts under .docpull/runs/<run_id>/, including run.json, documents.jsonl, chunks.jsonl, manifest.json, errors.jsonl, accounting.json, source-health.json, documents.ndjson, corpus.manifest.json, sources.md, and local.pack.json.
  • Add the project SQLite index for sources, runs, documents, chunks, errors, diffs, exports, and source health, with idempotent PRAGMA user_version schema setup.
  • Add deterministic project diffs with added, removed, changed, unchanged, title/path change, likely API behavior change, pricing change, and source health signals. Optional BYOK semantic summaries skip cleanly when no model key is configured.
  • Add agent context-pack exports for Cursor, Claude, Codex, OpenAI, LlamaIndex, and LangChain, plus stable JSONL eval-set generation from changed or latest documents.
  • Add project discovery persistence with docpull add URL --discover and docpull sync --update-discovery, including exact discovered URL refreshes.
  • Add a project diff demo launch asset at docs/launch-assets/docpull-project-diff-demo.png.

Changed

  • Reframe README and website copy around refreshable, cited, agent-ready context packs for changing public docs.
  • Clean project titles and dedupe alias pages before indexing/exporting run artifacts, reducing noisy duplicate source entries in real docs dogfood runs.

[5.0.2] - 2026-06-24

Added

  • Add docpull share for serving generated Markdown, HTML, or plain text reports at a simple local URL, with loopback-only binding by default and an explicit --allow-network-bind opt-in for exposed report links.

[5.0.1] - 2026-06-23

Added

  • Expose non-secret run accounting receipts in MCP structured responses for budget-blocked cloud/provider routes and local pack answers, including run.accounting.json links when durable artifacts exist.

[5.0.0] - 2026-06-22

Added

  • Add the local-first expansion surface: optional agent-browser rendering, provider-neutral discovery packs, source policy validation, pack refresh reports, pack audits, cited local answers, JSONL/agent exports, a localhost-only pack server, authenticated-source checks, and cron-friendly local monitors.
  • Add a shared free-first budget contract with CLI --budget, SDK BudgetConfig, policy-file budget.maximum_paid_cost_usd, route explanations, stricter effective paid caps, and deterministic run.accounting.json artifacts for budgeted or paid-capable runs.
  • Enforce fail-closed zero-dollar runs: local cache, direct HTTP, sitemap/static discovery, local extraction, local indexing, local pack intelligence, local monitors, and local agent-browser rendering remain available, while live Tavily, Exa, Parallel, Vercel Sandbox, E2B, provider probes, and paid-capable benchmark routes are blocked before execution.
  • Add the Phase 2 zero-dollar benchmark mode and target set, including completion classes for complete_for_0, complete_with_local_browser, partial_for_0, requires_provider, requires_cloud_browser, and blocked_by_policy.
  • Add provider-free discovery scanning with docpull discover scan URL for llms.txt, RSS/Atom feeds, OpenAPI references, richer sitemap discovery, and public GitHub documentation trees, all writing the standard candidate_sources.ndjson discovery-pack contract.
  • Add Phase 4 escalation suggestions when local capture is partial, with local discovery/render commands first, BYOK provider dry-run/live commands next, cloud rendering last, and estimated paid request/cost guards before escalation.
  • Add provider-neutral local parity workflows: docpull extract-pack, docpull map, docpull crawl-pack, docpull research-pack, and docpull entities-pack, plus SDK helpers and MCP tools over the same modules. These workflows write local lifecycle artifacts including events.ndjson, status.json, poll.report.json, and webhook.sample.json.
  • Add local structured-output validation for docpull research-pack --schema using a dependency-free JSON Schema subset over cited local answer fields.
  • Add monitor lifecycle controls: docpull monitor trigger, pause, unpause, and scheduler-snippet, plus monitor dedupe labels.
  • Add MCP tools for local expansion workflows: render_url, discover_sources, fetch_discovered_sources, extract_pack, map_sources, crawl_pack, research_pack, entities_pack, refresh_pack, audit_pack, answer_pack, validate_policy, export_pack, and serve_pack_status.
  • Add local downstream export formats for Sheets CSV/TSV, n8n workflow JSON, Vercel AI SDK JSON, CrewAI JSON, warehouse NDJSON, and optional Parquet via docpull[parquet].
  • Add explicit optional-renderer diagnostics: docpull render --check, docpull render --agent-browser-bin, SDK check_agent_browser_availability(), doctor reporting for the external agent-browser runtime, and a live smoke test that skips when the executable is absent.
  • Harden optional browser rendering by requiring HTTPS except localhost/loopback HTTP and rejecting non-default renderer action permissions.
  • Add optional cloud sandbox render runtimes: Vercel through the Vercel Sandbox CLI and E2B through the E2B Python SDK/API key.
  • Standardize cloud rendering on the same agent-browser --json contract as local rendering. Add docpull render --runtime local|vercel|e2b, docpull render init ..., docpull render doctor, estimated per-render budget caps, E2B template support, prebuilt sandbox install skipping, E2B file result transport, and opt-in live cloud smoke tests gated by DOCPULL_LIVE_CLOUD_RENDER=1.
  • Add local pack intelligence commands: docpull pack citations, docpull pack entities, docpull pack search, and docpull pack brief, plus matching MCP tools for agent access to citation maps, structured signals, cited pack search, and cited briefs.
  • Add docpull pack prepare, prepare_pack, and MCP pack_prepare to write the standard local pack intelligence bundle in one step.
  • Post-process successful Parallel, Tavily, and Exa provider context packs with local score, source-score, citation, entity, search, and brief artifacts.
  • Add first-class Tavily and Exa provider adapters plus docpull tavily ... and docpull exa ... aliases for context-pack and extract-pack workflows.
  • Add a provider capability matrix and docpull tavily map-pack, which uses Tavily Map to write a standard DocPull discovery pack.
  • Harden provider API key handling by rejecting unsafe key values before header use or local secret-file writes.
  • Move provider capability metadata out of adapter code, split provider tests by key/adapter/CLI responsibility, and add auth path redaction for CI/agent logs.
  • Add explicit provider live probes: Tavily account usage, Exa public team info, Parallel opt-in auth-gate validation, and guarded smoke probes separate from offline auth readiness.
  • Add make test-inventory and make test-all-local so contributors can report default pytest, fully gated pytest collection, Bun MCP, and coverage gates separately.

Changed

  • Reframe public docs around DocPull as a local-first, free-first evidence engine: local and open-data routes first, BYOK providers as explicit escalation, hosted execution as a future product boundary, and no hidden paid calls, CAPTCHA bypass, stealth scraping, or proprietary web-scale index claims.

Fixed

  • Resolve discovery links against the final redirected URL so moved docs sites do not produce broken relative links.
  • Prefer real streamed/hidden application content over loading skeletons when extracting modern docs pages, and remove loading-only placeholders from the selected content tree.
  • Clean up stdio MCP child processes in tests so the full local gate does not leak servers after client failures.

[4.4.1] - 2026-06-17

Changed

  • Remove the embedded README demo GIF from the PyPI long description so the project page stays focused on install and usage content.
  • Raise optional MCP dependency floors for starlette and cryptography so release audits install patched versions.

[4.4.0] - 2026-06-16

Added

  • Add release-ready SEC filing evidence packs with evidence.pack.json, AGENT_CONTEXT.md, validated rule confidence values, and a checked-in vendor-dependency rules example.

Changed

  • Wire docpull[proxy] to actual SOCKS proxy handling via aiohttp-socks, keep HTTP/HTTPS proxying on aiohttp's native request path, raise the optional proxy floor to aiohttp-socks>=0.11.0, and document security-floor dependencies as Renovate-managed constraints.
  • Run Renovate version checks across Python, GitHub Actions, web npm, and MCP Bun manifests so dependency floor updates are raised promptly.
  • Pin release build backends (setuptools, wheel) alongside pip, build, and twine, and run the publish build without isolation so releases do not silently download latest build tooling.

[4.3.1] - 2026-06-15

Changed

  • Tighten PyPI, GitHub, README, and website metadata around the public-web to agent-ready Markdown positioning.
  • Add launch copy, comparison guidance, and marketing visibility research for developer, Python, MCP, and RAG discovery channels.

[4.3.0] - 2026-06-14

Added

  • Add first-class Open Knowledge Format output via --format okf and --profile okf, including OKF concept frontmatter, generated directory index.md files with root okf_version: "0.1", and docpull corpus manifests.
  • Add the docpull.scraper API surface (Scraper, scrape_one, scrape_site, and ScrapeResult) as thin scraper-native names over the existing browser-free Fetcher pipeline.
  • Add docs/scraping-boundary.md to define docpull as a local, auditable static/server-rendered web-to-context scraper rather than a general browser automation framework.
  • Add SQLite FTS5 indexing for --format sqlite output plus search_sqlite_documents() for local full-text retrieval.
  • Add static Docusaurus, Sphinx, MkDocs/Material, VitePress, Starlight, GitBook, ReadMe.io, and Redoc/Scalar-style extraction fixtures so common docs frameworks are extracted and tagged without JavaScript rendering.

[4.2.0] - 2026-06-08

Added

  • Add docpull benchmark quick for repeatable real-site benchmark reports that compare core docpull crawls, cached reruns, and optional live Parallel Search / Search + Extract context-pack cases behind a local cost guard.
  • Add docpull benchmark article to turn benchmark JSON reports into a publishable Markdown draft with methodology, results, reproduce commands, and artifact links.
  • Add optional docpull[observability] Raindrop support so benchmark cases can be emitted as metadata-only traces when RAINDROP_WRITE_KEY is configured.
  • Add Raindrop event ids and per-case signals for benchmark failures, low scores, high-score cells, high-cost cells, and score-dimension warnings.
  • Add Tavily and Exa live benchmark cases that normalize provider results into the same scored context-pack artifacts as core docpull and Parallel.
  • Add docpull benchmark quick --target-set tool-docs/provider-matrix for provider-by-target matrix evals across Parallel, Exa, Tavily, Raindrop, DocPull, and low-cap adversarial public targets. The old v2 target-set name remains a compatibility alias.
  • Add weighted benchmark sub-scores for coverage, cleanliness, source fidelity, freshness, and density so clean and noisy targets no longer collapse to the same headline score.
  • Add Tavily credit-to-dollar normalization with --tavily-credit-usd or TAVILY_CREDIT_USD, plus a weekly GitHub Actions provider-matrix benchmark.
  • Add docpull providers for equal optional Parallel, Tavily, and Exa key status, durable key setup, and provider context-pack runs that can use any configured subset.
  • Add Make targets for quick, Parallel, and Raindrop benchmark runs.

Changed

  • Let docpull benchmark quick --provider auto/all run all locally configured providers and skip missing API keys or optional SDKs without failing the core benchmark.
  • Write sources.md for LLM-profile NDJSON output so core docpull packs score consistently with Parallel-generated context packs.
  • Score Parallel Search packs from their search metadata and keep fallback-pack core extraction artifacts scoped to their intended output directory.

[4.1.0] - 2026-06-07

Added

  • Add optional docpull[parallel] support for building Parallel Search + Extract context packs with local NDJSON, source Markdown, manifests, and workflow metadata.
  • Add docpull parallel import for offline fixture/demo workflows and a checked-in Parallel context-pack example fixture.
  • Add docpull parallel demo, backed by a packaged fixture, so the offline context-pack demo works from an installed wheel.
  • Add a Parallel product cross-reference covering Search, Extract, Task, FindAll, Entity Search, Monitor, MCP, and planned follow-up workflows.
  • Align Search mode choices with Parallel API docs (turbo, basic, and advanced) and request a Task text output schema for --task-brief.
  • Add source-policy, client-model, dry-run, and local cost-guard controls for live Parallel context packs.
  • Add broader Parallel artifact workflows for Entity Search, FindAll, TaskGroup batches, Monitor metadata/events, and llms.txt/OpenAPI API packs.
  • Add docpull parallel search-pack, extract-pack, task-pack, task-result, and task-events for Search-only, known-URL Extract, and Task lifecycle packs.
  • Add FindAll ingest, result, schema, enrich, extend, cancel, and events pack workflows.
  • Add snapshot monitor creation, monitor source-policy/location/webhook/metadata controls, event-group summaries, and checked-in Parallel API-pack recipes.
  • Add MCP tools for parallel_context_pack, parallel_api_pack, pack_score, and pack_diff, plus a built-in parallel source alias.
  • Add docpull pack score and docpull pack diff for local pack quality checks and refreshed-pack comparisons.
  • Add docpull parallel auth to check optional SDK and PARALLEL_API_KEY readiness without storing or printing secrets.
  • Add raw Markdown/plain-text conversion for docs indexes such as llms.txt through the normal fetch pipeline.
  • Add Parallel fetch policy, excerpt-size, and Search location controls to context-pack CLI and recipe workflows.
  • Add Monitor list, retrieve, update, cancel, trigger, and cursor/event-group events pack workflows.

Changed

  • Cap docpull parallel context-pack --extract-limit and context-pack recipes at 20 URLs so a single Parallel Extract request stays within the documented API limit.
  • Make docpull parallel taskgroup-pack --wait poll TaskGroup status until the group is inactive before snapshotting run outputs.
  • Treat no-content, invalid-content, HTTP-error, and save-empty skips as failures for docpull --single, so single-page agent fetches do not report empty output as success.
  • Extend docpull parallel run beyond context-pack recipes so YAML/JSON recipes can dispatch the same Parallel pack workflows as the explicit CLI commands.
  • Refresh the documented 10,000-page benchmark wall time to the latest local audit run.

Security

  • Route remote docpull parallel api-pack sources through docpull's hardened HTTPS-only URL validation, robots.txt check, DNS-pinned HTTP client, redirect revalidation, and response-size cap instead of a raw urllib fetch.

[4.0.1] - 2026-06-06

A release-readiness patch that tightens the public product boundary. No runtime API changes and no migration needed.

Changed

  • Make the Python docpull mcp server the only documented supported MCP path for agents, plugins, Claude Code, Cursor, and Claude Desktop.
  • Mark the root TypeScript/Bun mcp/ tree as an internal lab, make its package metadata private, and remove end-user install instructions for that path.
  • Replace stale YAML example files with current CLI recipes so docs no longer advertise removed options such as --sources-file, TOON output, keep_variant, language, or create_index.
  • Update website examples and performance copy to match the current CLI and benchmark results.

[4.0.0] - 2026-06-04

A security + cleanup release. A multi-agent security audit closed a high-severity SSRF and nine further findings (see Security); it ships alongside a tech-debt cleanup that removes several unused public APIs (see Removed — the breaking changes that make this a major release).

Security

  • DNS-rebinding TOCTOU in the URL validator (high). UrlValidator.resolve_allowed_addresses() resolved the hostname a second time and used that unscreened answer as the connect target, so a TTL-0 attacker could pass validation with a public IP and have the socket dialed at an internal one (e.g. cloud metadata). It now resolves once and returns exactly the addresses it screened.
  • Wider SSRF coverage. Block CGNAT shared address space (100.64.0.0/10) and IPv4-mapped IPv6 forms, and strip the trailing DNS root dot (localhost.) before the localhost/suffix checks — in both the Python validator and the TypeScript MCP source gate. The MCP gate additionally denies wildcard DNS-rebinding hosts (*.nip.io, *.sslip.io, *.xip.io).
  • robots.txt memory-exhaustion DoS. Cap the robots.txt body read at 512 KB, matching the existing sitemap limit.
  • YAML frontmatter injection. Frontmatter list items (tags/keywords sourced from page JSON-LD / OpenGraph) are quoted, escaped, and stripped of CR/LF so a hostile page cannot inject top-level frontmatter keys.
  • Conditional-request header injection. Cached ETag / Last-Modified values are stripped of CR/LF/NUL before being reused as If-None-Match / If-Modified-Since.
  • Supply chain. Pin release tooling (pip / build / twine) via requirements-release.txt; drop six unused (ghost) dependencies from the MCP package; bump aiohttp to >=3.14.0 (CVE-2026-34993, CVE-2026-47265).

Removed

  • Unused public methods on CacheManager (breaking for any external caller): has_changed, is_fetched, is_failed, get_failed_urls, get_cache_stats, clear_state, and has_resume_data had no callers in the library or tests. Incremental fetch and resume are unaffected — they use the retained update_cache, mark_fetched, mark_failed, get_fetched_urls, get_pending_urls, save_/load_/clear_discovered_urls, and evict_expired.
  • StreamingDeduplicator.is_duplicate (breaking): unused read-only probe. Use check_and_register, whose first return value reports whether content was new.
  • DocpullConfig.from_yaml_file (breaking): unused convenience wrapper. Use DocpullConfig.from_yaml(path.read_text()).

Changed

  • Internal cleanup with no API or behaviour change: removed dead code (the unused concurrency package, logging_config, and several private dead methods) and de-duplicated the discovery HTML-fetch helper and the HTTP GET/HEAD redirect re-validation path.

Fixed

  • MCP indexing of large libraries. pgvector embedding inserts are now batched under PostgreSQL's 32767 bind-parameter ceiling, so a library with thousands of chunks indexes in one transaction instead of failing.

[3.0.2] - 2026-05-29

A small release hygiene patch for the MCP hardening release. No API changes; no migration needed.

Fixed

  • Runtime version now matches package metadata. docpull --version and the default HTTP User-Agent report 3.0.2 instead of the stale 3.0.0 value that remained in docpull.__version__ after the 3.0.1 publish.

Tests

  • Added a stdio MCP smoke test. The test starts docpull mcp through the official MCP client, verifies the advertised 8-tool surface, checks structured list_sources output, and confirms SSRF rejection still flows through a real MCP call_tool request.

[3.0.1] - 2026-05-29

A security and correctness patch. No API changes; no migration needed.

Security

  • User-defined MCP sources are validated on load. Entries in ~/.config/docpull-mcp/sources.yaml are now rejected unless the name is a safe identifier, the URL is HTTPS to a public host (private, loopback, link-local, and internal-suffix hosts are blocked), and max_pages is in range. Previously a hand-edited config could point ensure_docs at an internal address.
  • grep_docs bounds regex execution per line. On top of the existing total wall-clock budget, each line now matches under a per-line timeout, closing the remaining catastrophic-backtracking (ReDoS) window for a pathological pattern against a single long line.

Fixed

  • Cache timestamps are timezone-aware UTC. Persisted timestamps (cache manifest, save steps, MCP metadata) use UTC consistently; legacy naive timestamps are parsed as UTC so cache-TTL comparisons stay deterministic instead of mis-expiring entries.
  • Swallowed exceptions are now logged. robots.txt parsing, the OpenAPI and SPA heuristics, and link extraction log skipped or invalid input at debug level instead of silently dropping it.

[3.0.0] - 2026-04-26

The deprecations 2.4 promised. Six config fields that have emitted a DeprecationWarning since 2.4 are now gone, and the naming_strategy literal no longer accepts the "flat" / "short" aliases that were documented as "aliased to 'full' until 3.0".

Breaking

  • ContentFilterConfig removed fieldslanguage, exclude_languages, deduplicate, max_total_size, exclude_sections. All have been no-ops since 2.4 with a deprecation warning on use; pydantic will now reject configs that set them (model_config = {"extra": "forbid"}). For deduplicate=True, switch to streaming_dedup=True. The other fields had no replacement because they had no effect.
  • OutputConfig.create_index removed — also a no-op since 2.4. Drop the field from your config; nothing to migrate.
  • OutputConfig.naming_strategy literal narrowed — the alias values "flat" and "short" (which silently behaved like "full") are no longer accepted. Use "full" directly. "hierarchical" is unchanged.
  • docpull.deprecated logger removed — the dedicated logger and the per-call DeprecationWarning infrastructure for the above fields are gone with them. Filters that targeted docpull.deprecated can be removed.

Migration

If your config file or DocpullConfig(...) call sets any of the removed fields, delete those lines. Pydantic's forbid policy will otherwise raise ValidationError at construction time with a clear "Extra inputs are not permitted" message naming the field.

[2.5.1] - 2026-04-25

A small but real bugfix: the grep_docsread_doc round-trip was broken. grep_docs returned paths with the library name prepended (e.g. hono/middleware/basic-auth.md), but read_doc joins library and path itself, so passing a grep result verbatim produced hono/hono/middleware/basic-auth.md and 404'd. The contract advertised in the read_doc description ("the natural follow-up to grep_docs: pass the library + path it returned") didn't actually work.

Fixed

  • grep_docs returns library-relative paths. Each result in the structured files payload now has both library (the library name) and path (relative to the library root). Pass them straight into read_doc(library=..., path=...) — no munging. Human-readable text rendering still shows library/path as the qualified identifier, so existing terminal output looks identical.
  • Tool descriptions for grep_docs and read_doc updated to match the actual contract.

Schema

  • _GREP_DOCS_OUTPUT_SCHEMA.files.items now requires library in addition to path. Existing consumers that read path will get a different (now correct) value; consumers that don't pipe grep results into read_doc are unaffected.

Tests

  • Added test_grep_to_read_doc_roundtrip and test_grep_to_read_doc_roundtrip_with_line_slice regression tests that pass library and path from grep verbatim into read_doc and assert success.
  • Added test_grep_docs_path_is_library_relative_in_subdir to cover nested files.
  • Updated test_grep_docs_structured_payload to assert the new library field and exact (not just suffix-matched) path value.

[2.5.0] - 2026-04-25

A focused MCP-server hardening pass. Closed three exploitable security holes in the agent-facing tools, added the missing ToolAnnotations that gate Anthropic Directory submission, exposed structured output alongside the rendered text on every tool that carries data, and added the three tools an agent obviously wants — read_doc to follow up a grep_docs hit, plus add_source / remove_source to manage the user registry programmatically.

Added

  • read_doc(library, path, line_start?, line_end?) — read a Markdown file from a fetched library, optionally line-sliced. The natural follow-up after grep_docs returns a hit; agents no longer need filesystem access for surrounding context. Path is resolved and confirmed to stay under the library root.
  • add_source(name, url, ...) — add or update a user source alias in the writable sources.yaml. Refuses to shadow a builtin alias unless force=true; URL is HTTPS-only and validated against the same SSRF rules as fetch_url. Atomic write (tmp + rename).
  • remove_source(name, delete_cache?) — remove a user source alias and optionally its cached Markdown directory. Cannot remove builtins (suggest add_source(force=true) to shadow instead). Cache deletion does a defense-in-depth resolved-path check.
  • ToolAnnotations on every toolreadOnlyHint / destructiveHint / idempotentHint / openWorldHint / title. Required for Anthropic Directory submission and unlocks host auto-approve for the four read-only tools.
  • Server instructions — system-prompt hint telling agents the call ordering (list_sources → ensure_docs → grep_docs → read_doc).
  • Progress notificationsensure_docs forwards FETCH_COMPLETED events as MCP progress to clients that supplied a progressToken on the call.
  • Structured output (outputSchema + structuredContent) on list_sources, list_indexed, grep_docs, read_doc, ensure_docs, add_source, remove_source. Clients that consume structuredContent get parseable JSON; clients that don't still see the rendered Markdown text.

Fixed

  • SSRF in fetch_url — schema previously accepted any string with no scheme/host enforcement. An agent could request http://169.254.169.254/, http://localhost, file:///etc/passwd, etc. Now validated upfront with the same UrlValidator (HTTPS-only, no localhost / private / link-local IPs) the crawler uses, instead of relying on the slow pipeline error path.
  • Path traversal in grep_docs / read_doc via librarydocs_dir / library did not validate library, so library="../../etc" walked anywhere the process could read. read_doc's path arg was similarly unchecked. Both now reject unsafe names (is_safe_library_name) and read_doc additionally resolves the joined path and confirms it stays under the library root.
  • ReDoS in grep_docs — pattern was compiled with no length cap and run line-by-line over every cached .md; Python re has no timeout knob. Now cap pattern length at 1000 chars and apply a 10s wall-clock budget across files.
  • isError flag was being silently dropped — the previous _call_tool returned a bare list[TextContent], which the SDK's legacy path hardcodes as isError=False regardless of what the handler intended. Every error your tools raised was being reported to clients as success. Now _call_tool returns CallToolResult directly so isError propagates correctly.
  • ensure_docs partial-fetch detection — a crash mid-crawl used to leave files on disk with no meta, and the next call would re-fetch (correct, but wasteful) or — if a stale meta from a prior run was present — trust the half-fetched cache. Meta writes are now atomic (tmp + rename) and a partial=true flag marks half-fetches so _cache_fresh treats them as stale.
  • grep_docs honors context > 1 — the schema advertised maximum: 3 but the implementation only ever rendered one line either side. Now renders up to context lines on each side.
  • load_user_sources silently swallowed YAML errors — a typo in the user's sources.yaml produced "Unknown source" instead of surfacing the parse failure. Now logs a warning at WARNING level.

Changed

  • Tighter input validation in the MCP _call_tool dispatcher: required strings checked with _require_str, ints coerced with _coerce_int. Errors that used to surface as ugly "invalid literal for int(...)" now return clear messages naming the bad argument.
  • Tighter input schemas: https:// pattern on fetch_url.url, enum on category, regex + maxLength on library everywhere it appears, maxLength: 1000 on grep_docs.pattern, integer bounds on max_tokens / max_pages.
  • _cache_fresh now also requires the source directory to contain at least one .md file — a manually-rm -rf'd cache no longer reports as fresh.
  • _PROFILE_ALIASES mapping deleted; _resolve_profile now goes through ProfileName directly, eliminating drift.
  • 39 new MCP tests (61 total in test_mcp_tools.py, 316 in the full suite). Coverage includes SSRF rejection, path traversal, oversized regex, partial-meta freshness, structured payloads, and the new write tools.

[2.4.0] - 2026-04-26

A two-pass cleanup. The first pass closed every claim the code didn't back (seven scaffolded-but-unwired config fields, the no-op fence-language regex, the "ETag-based caching" claim that never sent If-None-Match, the robots.txt UA mismatch, the MdxSourceExtractor dead branch). The second pass earned the local-first / agent-native / zero-trust pitch the marketing points at: cookie-banner stripping, rich frontmatter, --skill mode, streaming discovery, conditional GET, and a measured 10k-page benchmark.

Added

  • Conditional GET on cached pages: FetchStep now sends If-None-Match and If-Modified-Since from the manifest, and a 304 Not Modified response short-circuits with SkipReason.CACHE_UNCHANGED. Previously the marketing claimed ETag-based skipping but the headers were never actually sent. Re-runs against an unchanged site now transfer near-zero bytes.
  • Hierarchical naming: output.naming_strategy: hierarchical (set by the Mirror profile) preserves URL paths as nested directories (/api/auth/oauth2api/auth/oauth2.md), with sanitized segments and trailing-slash → index.md collapse. Path-traversal segments (..) are neutralized so URL-driven escapes can't leave the output dir.
  • --skill NAME generates agent-ready skills/rules: docpull URL --skill foo --skill-agent all writes scraped pages under .docpull/skills/foo/references, creates Claude Code and Codex SKILL.md wrappers, and writes a Cursor .cursor/rules/foo.mdc project rule. With --output-dir, the corpus is staged there while explicit agent targets still write active wrappers. The manifest description is derived from the first page's OpenGraph or JSON-LD metadata, with --skill-description available as an explicit override.
  • --require-pinned-dns refuses proxy configurations that delegate DNS to the proxy. Before this, running with --proxy silently weakened the SSRF posture (only a startup warning was logged); the new flag makes the trade-off explicit. Default off; intended for agent-driven workflows.
  • --no-streaming-discovery: backstop flag for the new producer- consumer fetch pipeline (see Changed). Falls back to the legacy discover-all-then-fetch behavior in case backpressure regressions surface in the wild.
  • max_file_size content-filter is now wired to the HTTP client's per-response cap. Previously hardcoded at 50 MiB; users can now lower it for OOM-prevention on runaway responses.
  • Rich frontmatter: every Markdown file now ships with a heading outline (top-level h1/h2, ≤12 entries), an ISO 8601 crawled_at timestamp, OpenGraph description, and a whitelisted slice of JSON-LD/microdata fields (author, published_time, keywords, etc.). Previously OG/JSON-LD extraction ran but the result was dropped.
  • MCP surface polish: ensure_docs accepts a profile argument (rag/mirror/quick/llm); grep_docs ranks results by per-file match density and renders ±1 line of context per hit (configurable via context); list_indexed reports humanized fetch age per source; fetch_url includes chunk count in its response header.
  • 10,000-page benchmark: tests/benchmarks/test_10k_pages.py stands up a synthetic localhost site with injected duplicates and reports wall time, peak RSS delta, manifest size, p50/p95/p99 per-page latency, and time-to-first-save. Gated behind DOCPULL_BENCHMARK_10K=1. README's new ## Performance section documents the headline numbers.

Fixed

  • Code-fence language normalization: html2text emits [code]…[/code] blocks without language tags by default, and the post-conversion regex meant to fix this was a self-replace no-op. Pages with Prism (class="language-python"), highlight.js (lang-py / hljs-language-X), Shiki, or GitHub-style (highlight-source-rust) syntax classes now produce GFM fenced blocks with the right language tag. plaintext / text / none are correctly treated as "no language."
  • Cookie / consent banner leakage: the FAQ claimed common banners were stripped, but no selectors targeted the vendor SDK shapes. Added selectors for OneTrust, Osano, Cookiebot, CookieLaw, CookieConsent, Iubenda, Termly, and generic .cookie-* / .gdpr-* / .consent-* patterns plus aria-label*="cookie|consent|gdpr" fallbacks. Pages that legitimately discuss cookies in their body are unaffected (selectors are structural, not text-based).
  • Streaming dedup hashed full Markdown including frontmatter: meant two URLs serving byte-identical body content never deduped because their source: and (after this release) crawled_at: fields differed. Dedup now strips frontmatter before hashing — the point of streaming dedup is "same body content," not "same bytes."
  • Robots.txt UA mismatch: docpull matched robots.txt rules as docpull/2.0 while sending requests as Mozilla/5.0 ... AppleWebKit .... Site operators scoping rules at User-Agent: docpull got no effect. Both surfaces now use the same UA, derived from the HTTP client.
  • Empty cache-fresh short-circuit when output is missing: if a user cleared output/ but kept cache/, the new conditional GET would have skipped on 304 and left no Markdown on disk. FetchStep now suppresses conditional headers when the expected output file is absent, forcing a fresh re-fetch.

Changed

  • Default User-Agent: docpull/{version} (+https://github.com/raintree-technology/docpull), replacing the previous Mozilla camouflage. Sites that whitelisted Mozilla patterns may need to allow the new UA; --user-agent continues to override. The pitch is "polite crawler" — disguising as a browser contradicted that.
  • Streaming discovery → fetch (default): URLs now flow through a bounded worker pool as the discoverer yields them, instead of being collected to a list before any fetching begins. First PAGE_SAVED on a 10,000-page synthetic site now fires within ~70 ms of the run starting (vs. waiting for full discovery before). The discoverer awaits when url_queue is full, so backpressure self-regulates. --no-streaming-discovery falls back to the legacy path.
  • Mirror profile keeps flat naming by default in 2.x to preserve existing users' output paths. Hierarchical is now opt-in via --naming-strategy hierarchical (CLI) or output.naming_strategy (YAML). The Mirror profile default flips to hierarchical in 3.0; the upgrade will be flagged in the 3.0 release notes.
  • MdxSourceExtractor removed from DEFAULT_CHAIN (it always returned None). The class and find_mdx_source_url helper are still exported for callers that want to wire prefer_source manually.

Deprecated

The following config fields warn at runtime when set to a non-default value and will be removed in 3.0. Each was scaffolded but never read by the pipeline; CLI flags backed by them have been dropped from --help.

  • content_filter.language and --language (no language detector ever shipped — pursue if a real user asks)
  • content_filter.exclude_languages
  • content_filter.deduplicate (post-processing dedup; streaming_dedup covers the use case)
  • content_filter.exclude_sections
  • content_filter.max_total_size (cumulative byte budget across an async fetch is racy; per-page max_file_size is the right knob)
  • output.create_index (no INDEX.md generator; downstream tools don't need one)

Security

  • Honest UA disclosure: see Changed. Polite-crawling claim now lines up with what site operators see in their logs.
  • --require-pinned-dns: see Added. Closes a previously-undisclosed gap where --proxy disabled docpull's connector-level DNS pinning.
  • Hierarchical-naming traversal: URL path segments are sanitized (..index, special chars → _, runs of underscores collapsed) before being joined into the output path. The SaveStep base-dir guard remains the second line of defense.

[2.3.0] - 2026-04-24

Sharpened positioning around the agent / RAG use case, plus real bug fixes surfaced by validation against Next.js, Supabase, Anthropic, FastAPI, Tailwind, and Drizzle documentation sites.

Added

  • Framework-specific fast extractors: Next.js __NEXT_DATA__, Mintlify, OpenAPI / Swagger JSON rendered directly to Markdown, plus source-type tagging for Docusaurus and Sphinx. Runs before the generic extractor.
  • Next.js App Router detection via self.__next_f.push, router state tree, and /_next/static/ path markers — no longer relies on __NEXT_DATA__, which is absent on modern App Router pages.
  • SPA detection (pre- and post-conversion): pages that produce only Loading... shells are skipped with a clear reason. --strict-js-required turns this into a hard error for agents that want to route elsewhere.
  • Trafilatura extractor as an optional alternative content extractor (pip install docpull[trafilatura], then --extractor trafilatura).
  • Token-aware Markdown chunking: --max-tokens-per-file N splits pages on heading then paragraph boundaries. Exact counts with tiktoken, character-estimate fallback otherwise.
  • NDJSON output format (--format ndjson) for streaming one record per page or per chunk. --stream writes to stdout for live pipeline consumption.
  • llm profile: bundles NDJSON + 4k-token chunks + rich metadata + dedup.
  • --single / fetch_one(url): fast single-page path with no discovery, designed for AI-agent tool loops.
  • Python MCP server (docpull mcp): exposes fetch_url, ensure_docs, list_sources, list_indexed, and grep_docs tools over stdio. Install via pip install docpull[mcp].

Fixed

  • robots.txt redirect handling: Cloudflare/HTTP-2 responses send lowercase header names, but the Location lookup was case-sensitive, causing 301/308 redirects to be treated as errors. This blocked docs.anthropic.com and any other site whose robots.txt was redirected.
  • html2text link escape artifacts: cleaned up mangled Markdown link targets containing prefix/<https:/real.url> in the post-processing pass; handles both text and image-only (empty-text) links.

Removed

  • Dead dependencies: requests (replaced by aiohttp in v2.0) and gitpython (never used in v2+).

Changed

  • ContentFilterConfig gains extractor, enable_special_cases, and strict_js_required fields. OutputConfig gains max_tokens_per_file, tokenizer, emit_chunks, and ndjson_filename.

2.0.0 - 2025-11-29

Breaking Changes

  • Complete architecture rewrite with new Python API
  • Moved to src/ layout (PEP 517/518 compliant)
  • Old GenericAsyncFetcher replaced by Fetcher class with async context manager
  • Configuration now uses Pydantic models (DocpullConfig)

Added

  • Streaming event API: Async iterator interface for real-time progress tracking
  • CacheManager: Persistent caching with O(1) lookups, batched writes, TTL eviction
  • StreamingDeduplicator: Real-time duplicate detection during fetch
  • Profiles: Built-in rag, mirror, quick profiles with sensible defaults
  • CLI cache options: --cache, --cache-dir, --cache-ttl, --no-skip-unchanged
  • Pipeline architecture: Modular steps (Validate, Fetch, Convert, Dedup, Save)

Changed

  • Cache uses sets internally for O(1) URL membership checks
  • Consistent SHA-256 hashing across cache and dedup (accepts str or bytes)
  • ETag and Last-Modified headers now extracted and cached

Removed

  • Old fetcher classes (GenericAsyncFetcher, AsyncDocFetcher, etc.)
  • DedupTracker replaced by StreamingDeduplicator
  • Legacy config fields (incremental, update_only_changed)

1.5.0 - 2025-11-28

Added

  • Proxy support: HTTP, HTTPS, SOCKS5 via --proxy or DOCPULL_PROXY env var
  • Retry with exponential backoff: --max-retries, --retry-base-delay for transient failures
  • Better encoding detection: Intelligent charset detection for international docs
  • URL normalization: Reduces duplicate fetches by 10-20%
  • Content hash change detection: SHA-256 hashing for efficient incremental updates
  • Custom User-Agent: --user-agent flag

Changed

  • robots.txt compliance is now mandatory (cannot be disabled)
  • Automatically respects Crawl-delay directives

1.4.0 - 2025-11-28

Breaking Changes

  • Removed profile system entirely - use URLs directly
  • --source flag removed; use positional URL arguments
  • Python API: url parameter instead of url_or_profile

1.3.0 - 2025-11-20

Added

  • Rich metadata extraction: --rich-metadata extracts Open Graph, JSON-LD, microdata
  • Enhanced frontmatter with author, description, keywords, images, publish dates

Changed

  • Removed 7 built-in profiles; generic fetcher works for all sites

1.2.0 - 2025-11-16

Added

15 major features for optimization and workflow automation:

Optimization

  • --language / --exclude-languages: Filter by language
  • --deduplicate / --keep-variant: Remove duplicate files
  • --max-file-size / --max-total-size: Size limits
  • --exclude-sections: Remove verbose sections

Output

  • --format: markdown, toon, json, sqlite
  • --naming-strategy: full, short, flat, hierarchical
  • --create-index: Generate INDEX.md

Workflow

  • --sources-file: Multi-source YAML configuration
  • --incremental / --update-only-changed: Resume and update detection
  • --git-commit / --git-message: Git integration
  • --archive / --archive-format: Create archives
  • --post-process-hook: Python plugin system

Changed

  • PyYAML and GitPython now required dependencies

1.1.0 - 2025-11-14

Added

  • --doctor command for installation diagnostics
  • TROUBLESHOOTING.md documentation

1.0.0 - 2025-11-07

Added

  • Initial release
  • Async + parallel fetching
  • Security: HTTPS-only, path traversal protection, XXE protection, size limits
  • Rate limiting and timeout controls
  • YAML frontmatter in output files