Proxy overview
August 2, 2026 · View on GitHub
TRAWL can expose a general HTTP/HTTPS forward proxy, normally on port 8192. It behaves like a
direct proxy for ordinary traffic and invokes the existing /scrape tier pipeline only when a
small buffered response looks like a supported challenge.
The proxy is disabled by default. It is intended for trusted clients on localhost, a private LAN, or a private container network.
Traffic flow
Client
│
├─ HTTP request ─────────────────────────────┐
│ │
└─ HTTPS CONNECT → local TLS termination ────┤
▼
Tier 0 direct forward
│
┌────────────────────────┼────────────────────────┐
│ │ │
normal small response challenge detected large/media response
│ │ │
buffer + return run scrape tiers stream
│
Tier 1 → Tier 2 → Tier 3 → Tier 4
│
return solved HTML or raw bytes
HTTPS requires local TLS termination so TRAWL can inspect the response body. TRAWL generates a per-host certificate signed by its own root CA. The client must trust that root certificate.
Direct forwarding
Tier 0 opens a normal TCP or TLS connection to the destination and forwards the request method, body, and end-to-end headers. It preserves headers such as:
Authorization,Cookie,User-Agent,Referer, andOrigin- custom API headers
Range,If-Range, cache validators, and content metadata
Hop-by-hop headers are rebuilt or removed at the proxy boundary.
Responses not selected by the streaming policy—including small HTML, JSON, XML, text, and unknown
small payloads—are buffered. Compressed gzip, deflate, and Brotli responses are decoded only
for challenge inspection; the original bytes remain the response body. Chunked buffered responses
are de-chunked before being returned.
Challenge escalation
When Tier 0 detects a supported challenge wall, the request is passed to the same scrape()
orchestrator used by POST /scrape:
- plain HTTP fetch;
- cached browser session;
- fresh browser challenge solve;
- residential proxy solve, when configured.
For browser-tier HTML responses, the proxy returns the rendered solved DOM rather than the original challenge response. Binary responses use the raw response bytes when the winning tier exposes them.
A hostname that produced a supported challenge wall is cached for five minutes. During that window, later requests for the hostname skip Tier 0 and go directly to the tiered solver.
Challenge and CAPTCHA solving is best effort. A site can still reject the browser, require user interaction, bind clearance to an unsupported signal, or change its challenge implementation.
Buffering and streaming
The proxy buffers ordinary responses by default so challenge detection sees the complete body. It streams when at least one of these applies:
Content-Lengthis at least 8 MiB;- the content type is
video/*oraudio/*; - an unknown-length response has a known binary content type;
- the URL has a known media, archive, installer, disk-image, PDF, or font extension.
Streaming keeps large files out of TRAWL's memory. Streamed responses bypass body-based challenge detection, so the policy intentionally targets content that is unlikely to be an HTML challenge.
Range requests are forwarded unchanged. A compliant upstream 206 Partial Content response,
including Content-Range and Accept-Ranges, is passed back to the client. If the Range response
is instead a detectable challenge page, it can escalate through the normal solver pipeline.
WebSockets
HTTP and HTTPS WebSocket upgrade requests use a bidirectional byte relay after the upstream
101 Switching Protocols response. WebSocket frames are not buffered or interpreted.
The WebSocket handshake does not escalate through the browser solver. If a WebSocket endpoint requires a clearance cookie, the client must already have suitable credentials or obtain them through a preceding solved HTTP flow.
Current protocol boundaries
The current listener is an HTTP/1.1 proxy:
- one request is served per proxied TLS connection;
- HTTP keep-alive and pipelining are not reused;
- request bodies use
Content-Length; chunked uploads are not decoded; - duplicate upstream response headers are represented by the first observed value;
- proxy authentication is not implemented;
- HTTP/2 and HTTP/3 are not terminated between the client and TRAWL.
On a challenged request, the browser fallback converts a request body to UTF-8 because the public scrape contract currently accepts a string body. Ordinary direct forwarding preserves binary request bodies exactly.
Proxy versus API endpoints
| Interface | Best for | Response handling |
|---|---|---|
POST /v1 | FlareSolverr-compatible integrations | Returns HTML, cookies, and user agent |
POST /scrape | Native programmatic scraping | Returns tier metadata and solved content |
HTTP proxy :8192 | Applications that perform their own requests | Forwards HTTP traffic and escalates detected challenges |
Use /v1 or /scrape when the application explicitly supports a solver API. Use the proxy when the
application only knows how to make ordinary HTTP requests or when cookie handoff is not sufficient.