Web Scraper

May 9, 2026 ยท View on GitHub

A purpose-built fetch tool, separate from generic http_request / curl. It exists because the agent doesn't want raw HTML - it wants the article.

What it does

  • Fetches a URL.
  • Strips boilerplate (nav, ads, footer, scripts).
  • Returns clean text the agent can reason over.

Guardrails

  • Caps response at 1 MB - large pages get truncated, not silently dropped.
  • 20-second timeout - slow servers don't stall the conversation.
  • Subject to the same proxy and URL-guard rules as other network tools.

What it's good for

  • Reading articles, blog posts, docs pages, GitHub READMEs without the noise.
  • Following up on a Web Search result.
  • Summarising a single page on demand.

See also