AI Crawl Control + Pay Per Crawl

August 22, 2026 ยท View on GitHub

Last modified: 2026-08-21

GPTBot receiving a 402 challenge, then the article after presenting a Crawler-Payment token

Each token redeems exactly once (config).

The ai_crawl_control policy implements the "Pay Per Crawl" pattern: AI crawlers that arrive without a valid Crawler-Payment token receive 402 Payment Required along with a JSON challenge body. A crawler that wants the content reads the challenge, posts a payment to your billing system, and retries with the issued token in the Crawler-Payment header. Each token redeems exactly once.

The implementation ships an in-memory ledger seeded from config and an HTTPS-only HTTP ledger client for production.

Challenge body

The proxy emits two challenge shapes:

  1. Single-rail (default). A 402 with the Crawler-Payment header and a flat JSON body describing the price. This is the path legacy crawlers see.
  2. Multi-rail (opt-in). When the agent sends Accept-Payment: or one of the multi-rail Accept MIME types (application/sbproxy-multi-rail+json, application/x402+json, application/mpp+json), the proxy emits a 402 with Content-Type: application/sbproxy-multi-rail+json and a body that lists one entry per rail the operator declared (x402, MPP, Lightning), each with its own quote-token JWS.

This policy decides which requests are payable and what they cost. It does not settle payments. Settlement is a separate block, proxy.payments, which owns the rails, the durable intent store, and the rule that a paid request reaches the origin only after a committed record says the payment settled. See payment-settlement.md for that configuration and 402-challenge.md for the exact bytes of every challenge, credential, error, and receipt.

With proxy.payments configured, the settlement gate takes over this policy's 402s: challenges are compiled from the matched tier's price into signed payment requirements, persisted as durable intents before the 402 is written, and rendered in the selected rail's wire shape, with the signed quote token in this policy's configured header. The in-memory valid_tokens ledger and the HTTP ledger below then no longer redeem those requests, and the legacy rails: and quote_token: blocks on this policy should be dropped in favor of the rails under proxy.payments. The request-path sequence, the per-rail challenge shapes, and the failure posture live in payment-settlement.md.

Request flow

crawler GET /article
        User-Agent: GPTBot/1.0
proxy   <- 402 Payment Required
        Crawler-Payment: Crawler-Payment realm="ai-crawl" currency="USD" price="0.001000"
        Content-Type: application/json
        body: {"error":"payment_required","price":"0.001000","amount_micros":1000,"currency":"USD","target":"blog.example.com/article","header":"crawler-payment"}

crawler GET /article (after paying out-of-band)
        User-Agent: GPTBot/1.0
        crawler-payment: tok_a89be2...
proxy   <- 200 OK
        body: <article>

crawler GET /article (replay attempt)
        User-Agent: GPTBot/1.0
        crawler-payment: tok_a89be2...
proxy   <- 402 (single-use ledger; token already spent)

Configuration

policies:
  - type: ai_crawl_control
    price: 0.001
    currency: USD
    header: crawler-payment           # default
    crawler_user_agents:              # case-insensitive substring match
      - GPTBot
      - ChatGPT-User
      - ClaudeBot
      - anthropic-ai
      - Google-Extended
      - PerplexityBot
      - CCBot
    valid_tokens:                      # in-memory ledger
      - tok_a89be2f1
      - tok_b7cf012e
      - tok_c34f9a82
FieldTypeDefaultDescription
pricefloatunsetPrice emitted in the challenge body and the price= parameter of the challenge header. Used as the fallback when no tier matches.
currencystringUSDISO-4217 code surfaced in the challenge header and body.
headerstringcrawler-paymentHeader the crawler reads from the 402 response and writes to its retry.
crawler_user_agentslistcovers GPTBot, ChatGPT-User, ClaudeBot, anthropic-ai, Google-Extended, PerplexityBot, CCBot, FacebookBotCase-insensitive substring matches against the request User-Agent. Empty list treats every GET / HEAD as a crawler.
valid_tokenslist[]Seeds the in-memory ledger. Each token redeems once, then leaves the set.
tierslist[]Pricing tiers. First match wins. See "Tiered pricing" below.
ledgerblockunsetHTTP ledger client config. See "HTTP ledger" below. Mutually exclusive with valid_tokens.

Only GET and HEAD requests are subject to charging today. POST, PUT, PATCH, and DELETE pass through without charge.

Calling it

The runnable configuration is examples/ai-crawl-control/, which is the block above in front of a proxy origin, seeded with three single-use tokens (token-aaa-001, token-aaa-002, token-aaa-003). Start it:

make run CONFIG=examples/ai-crawl-control/sb.yml

Arrive as a known crawler with no token:

curl -sS -i -H 'Host: blog.local' \
  -H 'User-Agent: GPTBot/1.0' \
  http://127.0.0.1:8080/article
HTTP/1.1 402 Payment Required
content-type: application/json
crawler-payment: Crawler-Payment realm="ai-crawl" currency="USD" price="0.001000"
content-length: 142

{"error":"payment_required","price":"0.001000","amount_micros":1000,"currency":"USD","target":"blog.local/article","header":"crawler-payment"}

Two details in there are easy to get wrong. price is a string, not a JSON number, and it is rendered at six decimal places, so 0.001 in the config reads back as "0.001000". amount_micros carries the same figure as an integer count of millionths, which is the field to compute against, since it avoids parsing a decimal string into a float. target is the host and path being charged for, so a crawler can tell which resource the quote covers.

The response header name is whatever header: is set to, here crawler-payment. Its value repeats the canonical Crawler-Payment scheme name followed by the realm, currency, and price.

Now pay. Retry with a seeded token in that same header:

curl -sS -i -H 'Host: blog.local' \
  -H 'User-Agent: GPTBot/1.0' \
  -H 'crawler-payment: token-aaa-001' \
  http://127.0.0.1:8080/article

The paywall opens and the upstream's real response comes back:

HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
link: </licenses.xml>; rel="license"
Transfer-Encoding: chunked

The link header and the chunked encoding are automatic on any origin that has this policy: see rsl.md for what else the proxy does to advertise the license document.

Send that exact request a second time and it is refused, because the ledger is single-use and the token left the set when it was redeemed:

HTTP/1.1 402 Payment Required
crawler-payment: Crawler-Payment realm="ai-crawl" currency="USD" price="0.001000"

{"error":"payment_required","price":"0.001000","amount_micros":1000,"currency":"USD","target":"blog.local/article","header":"crawler-payment"}

That replay refusal is the property worth checking in your own run. A token that keeps working is a ledger that is not spending, and with the in-memory ledger it also means a proxy restart has reseeded valid_tokens from the config.

Finally, the same URL with no crawler User-Agent is not charged at all:

curl -s -o /dev/null -w '%{http_code}\n' \
  -H 'Host: blog.local' http://127.0.0.1:8080/article
# 200

The challenge fires on the User-Agent substring match, so ordinary browser and client traffic reaches the origin untouched.

Tiered pricing

a crawler hitting a charged article URL and getting 402 while the free preview path returns 200

Tiers price full HTML, Markdown feeds, and previews differently (config).

A flat per-site price is the right starting point but not the right long-term shape. Different routes carry different commercial value, and the same article in three formats (HTML, Markdown, PDF) is worth three different prices to a training crawler. The tiers: field lets you price by route pattern and content shape without forking the policy.

policies:
  - type: ai_crawl_control
    price: 0.0005                      # fallback when no tier matches
    currency: USD
    tiers:
      - route_pattern: /premium/*
        price:
          amount_micros: 5000          # \$0.005 per crawl
          currency: USD
        free_preview_bytes: 1024       # cooperative crawlers get 1 KiB free
        paywall_position: top_of_page
      - route_pattern: /articles/*
        price:
          amount_micros: 1000          # \$0.001 per crawl
          currency: USD
        content_shape: markdown        # Markdown form only
        free_preview_bytes: 4096
        paywall_position: inline
      - route_pattern: /articles/*
        price:
          amount_micros: 500           # \$0.0005 per crawl
          currency: USD
        content_shape: html
      - route_pattern: /docs/*
        price:
          amount_micros: 250
          currency: USD
FieldTypeDescription
route_patternstringPath matcher. Supports literal paths (/about) and a * suffix wildcard (/articles/*). First match wins; later tiers act as fallbacks.
price.amount_microsu64Price in micros (1e-6 of one unit of currency). 1000 micros = $0.001. Floats never enter the wire format.
price.currencystringISO-4217 code. Must match the policy-level currency for now.
content_shapeenumOne of html, markdown, json, pdf, other. Advisory; surfaced in metrics and the redeem payload but not yet used as a tier filter.
free_preview_bytesu64, optionalByte budget the crawler may read without paying. Surfaced in the challenge body so cooperative crawlers can decide up front whether the preview alone meets their need.
paywall_positionenum, optionalHint to the crawler about where the paywall sits in the rendered response: top_of_page (paywall replaces the entire body; any free preview is a separate excerpt), inline (free preview served inline, paywall follows in the same body), bottom_of_page (paywall after the full or near-full preview; discouraged for high-value content).

The first tier whose route_pattern matches wins. When no tier matches, the policy falls back to the top-level price and currency. An empty tiers list keeps the original flat-price behavior.

Per-shape pricing

content_shape is advisory: configurations may set the field on a tier so metrics and the redeem payload carry the shape, but the policy does not yet match against it. The wire format is stable, so configurations that set content_shape today will keep working when the resolver lands.

HTTP ledger

The in-memory ledger (valid_tokens:) is fine for tests, fixed-token issuance, or one-off content gates. Production deployments with multiple proxy replicas need a network-callable ledger so one token spends across all nodes. The HTTP ledger client speaks a JSON-over-HTTPS protocol with HMAC-SHA256 envelope signatures over a fixed eight-line canonical form.

policies:
  - type: ai_crawl_control
    price: 0.001
    currency: USD
    ledger:
      url: "https://ledger.internal"   # required; plain http:// is rejected
      trust_roots: []                  # optional PEM CA bundles for private PKI
      key_id: "sb-ledger-2026-q2"
      secret_ref:
        env: SBPROXY_LEDGER_HMAC_KEY   # env var holding the hex-encoded HMAC key
      workspace_id: "default"          # default: "default"
      idempotency_key_header: "Idempotency-Key"   # default
      timeout_ms: 5000                 # per-attempt timeout; default 5000
      retry:
        max_attempts: 5                # 1..=5; hard-clamped by the client
        initial_backoff_ms: 250        # default 250
        max_backoff_ms: 5000           # default 5000
      breaker:
        failure_threshold: 10          # consecutive failures that open; default 10
        success_threshold: 1           # half-open successes to close; default 1
        open_duration_ms: 5000         # default 5000

The HMAC key resolves through secret_ref, which takes either env: <VAR> (an environment variable holding the hex-encoded key) or secret: <name> (a logical secret resolved through the secrets layer). For dev configs and tests only, an inline key_hex: is honored when secret_ref is absent; it should not appear in a production sb.yml. The agent identity fields on the redeem payload come from the request-time agent-class resolver, not from ledger config.

For a ledger signed by a private CA, each trust_roots entry may contain one or more PEM CERTIFICATE blocks. These certificates are added to the system trust store rather than replacing it, so public HTTPS endpoints continue to validate normally. Malformed or empty PEM bundles fail at config load.

The client refuses to construct against a non-HTTPS url at config-load time. Plain HTTP is a hard error because the request envelope carries an HMAC over the body, and TLS is the only thing keeping the body itself confidential. A ledger: block also fails config load when the binary was built without the http-ledger feature; it never falls back silently to valid_tokens.

Request envelope

Every redeem call carries the eight-line canonical envelope:

{
  "v": 1,
  "request_id": "01HZX...",
  "timestamp": "2026-04-30T12:34:56.789Z",
  "nonce": "8f4a...32-hex...",
  "agent_id": "openai-gptbot",
  "agent_vendor": "OpenAI",
  "workspace_id": "default",
  "payload": {
    "token": "tok_abc...",
    "host": "blog.example.com",
    "path": "/articles/foo",
    "amount_micros": 1000,
    "currency": "USD",
    "content_shape": "markdown"
  }
}

The signature is HMAC-SHA256 over the canonical signing string (eight \n-separated fields, last one being the SHA-256 of the request body). The signature lands in the X-Sb-Ledger-Signature: v1=<hex> header. The v1= prefix reserves room for future MAC migrations without breaking peers.

Idempotency

Every attempt carries an Idempotency-Key header (a fresh ULID per logical operation). Retries reuse the same key; the ledger short-circuits the second attempt with the cached response. A different body under the same key returns 409 ledger.idempotency_conflict, which protects against accidental key reuse across operations.

Idempotency-Key is distinct from the envelope's request_id: the request id identifies the inbound 402 from the agent, while the idempotency key identifies a single conversation with the ledger about that request.

Retry and circuit breaker

Exponential backoff with full jitter, max 5 attempts, per-attempt deadline 5 s, total deadline 30 s. The base schedule is 0 ms, 250 ms, 500 ms, 1 s, 2 s, each with [0, base) jitter added. Retries fire only on:

  • network errors (DNS, TCP RST, TLS handshake, read timeout)
  • HTTP 429 (with Retry-After honored)
  • HTTP 502 / 503 / 504
  • error envelopes with retryable: true

Hard failures (ledger.token_already_spent, ledger.signature_invalid, ledger.bad_request) translate directly to a 402 to the crawler. There is no point retrying a token the ledger already rejected as spent.

The circuit breaker opens after 10 consecutive failures, half-opens after 5 s and admits one probe at a time, and closes on probe success. While a probe is outstanding, other redeem calls are refused the same way they are while the breaker is open, so a recovering ledger sees one request rather than the full crawl. While the breaker is open, the client returns a synthetic ledger.unavailable error without making the network call. The policy treats that as "ledger is down" and fails closed: the crawler gets a 503 with a ledger_unavailable JSON body and a Retry-After header. There is no on_ledger_failure knob; fail-closed is the fixed behavior, because failing open would hand out the content the paywall exists to price.

A ledger Retry-After propagates straight to the crawler on that 503 (defaulting to 5 seconds when the ledger did not send one), so the crawler knows when to come back.

Failure modes

Ledger responsePolicy action
200 success, redeemedPass the request through.
200 success, not redeemed402 with the challenge body. The token was valid format but the ledger refused (out of balance, expired).
409 token_already_spent402, no retry.
4xx other402, no retry, log at WARN.
5xx, transient envelope, breaker openFail closed: 503 with a ledger_unavailable body and Retry-After. Not configurable.

Agent classes and per-vendor pricing

An agent_class taxonomy lets metrics, audit logs, and ledger payloads attribute revenue per vendor. The agent class is resolved at request time via three signals (in order of confidence):

  1. Verified Web Bot Auth keyid matches an expected_keyids entry. Highest confidence.
  2. Forward-confirmed reverse-DNS suffix matches an expected_reverse_dns_suffixes entry. Strong confidence.
  3. User-Agent regex match. Advisory unless the policy explicitly trusts UAs.

Three reserved sentinels round out the resolver:

  • human is emitted when no automated-agent signal is present.
  • unknown is the fall-through bucket for an automated UA without a registry match.
  • anonymous is emitted for anonymous Web Bot Auth requests with no known keyid.

Operators see all three values in metrics and dashboards; alerting on a sustained climb in unknown is the normal way to spot a new crawler that needs a registry entry.

Per-vendor pricing example

agent_classes:
  catalog: inline
  entries:
    - id: openai-gptbot
      vendor: OpenAI
      purpose: training
      expected_user_agent_pattern: "(?i)\\bGPTBot/\\d"
      expected_reverse_dns_suffixes: [".gptbot.openai.com"]
      expected_keyids: ["openai-2026-01"]
    - id: anthropic-claudebot
      vendor: Anthropic
      purpose: training
      expected_user_agent_pattern: "(?i)\\bClaudeBot/\\d"
      expected_keyids: ["anthropic-2026-01"]
    - id: commoncrawl-ccbot
      vendor: Common Crawl
      purpose: archival
      expected_user_agent_pattern: "(?i)\\bCCBot/\\d"

policies:
  - type: ai_crawl_control
    currency: USD
    tiers:
      # Training crawlers pay full price.
      - route_pattern: /articles/*
        agent_id: openai-gptbot
        price: { amount_micros: 2000, currency: USD }
      - route_pattern: /articles/*
        agent_id: anthropic-claudebot
        price: { amount_micros: 2000, currency: USD }
      # Archival crawlers get a discount.
      - route_pattern: /articles/*
        agent_id: commoncrawl-ccbot
        price: { amount_micros: 500, currency: USD }
      # Sentinel buckets price differently for diagnostics.
      - route_pattern: /articles/*
        agent_id: anonymous
        price: { amount_micros: 1000, currency: USD }
      - route_pattern: /articles/*
        agent_id: unknown
        price: { amount_micros: 1500, currency: USD }

agent_id on a tier matches against the resolver's verdict. The first tier whose route pattern AND agent id both match wins. A tier without agent_id matches every agent. expected_keyids lets a verified Web Bot Auth signature classify the request even when the User-Agent string is missing or spoofed.

Note the shape of every expected_user_agent_pattern above, because the proxy compiles the pattern exactly as you wrote it and adds nothing. The leading (?i) is yours: without it the pattern is case-sensitive, gptbot/2 does not match GPTBot/\d, and the request is priced as unknown instead. The proxy warns once per such entry at load. The \b and the /\d are yours too: the match is a substring search, not an anchored one, which is what makes a pattern work against a compound header like Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot). A bare MyPartnerBot with no delimiter therefore also matches Mozilla/5.0 (compatible; MyPartnerBot-imposter) and hands that client your partner's tier, so put a boundary the impersonator cannot append after your name. A User-Agent is unauthenticated in both directions in any case: expected_keyids and expected_reverse_dns_suffixes are what verify a crawler, and the UA pattern only classifies one.

The default agent classes ship embedded in the binary. Use catalog: inline when you want sb.yml to provide a complete catalog with your own expected_keyids; use catalog: builtin or omit the block to keep the embedded catalog.

Observability

Every redeem fires a metric and a structured-log line. The label set:

LabelSourceCardinality cap
agent_idAgent-class resolver. Bounded to registry plus human, unknown, anonymous sentinels.200
agent_classClosed enum from the taxonomy.8
agent_vendorFree-form vendor name from the taxonomy.20
payment_railClosed enum: none, x402, mpp_card, mpp_stablecoin, stripe_fiat, lightning.6
content_shapeClosed enum: html, markdown, json, pdf, other.5

Cardinality budgets are enforced by sbproxy-observe::cardinality::CardinalityLimiter; over-cap label values demote to __other__ and increment sbproxy_label_cardinality_overflow_total.

Metrics

MetricTypeNotes
sbproxy_ledger_redeem_duration_seconds{host, outcome}histogramLatency of the ledger round-trip, one observation per redeem. outcome is success, transient_failure, or hard_failure; count redeems from the histogram's _count series. Carries trace exemplars. There is no separate sbproxy_ledger_redeem_total counter.
sbproxy_circuit_breaker_transitions_total{origin, from_state, to_state}counterBreaker flap counter, shared with every other circuit breaker in the proxy. There is no ledger-specific breaker-state gauge.
sbproxy_requests_total{hostname, method, status, agent_id, agent_class, agent_vendor, payment_rail, content_shape}counterPer-request outcome.

The per-agent dashboard (deploy/dashboards/per-agent.json) groups every panel by agent_class plus agent_vendor, so operators see one row per vendor and one row each for the sentinels. The audit-log dashboard (deploy/dashboards/audit-log.json) shows admin actions on ai_crawl_control tier edits.

Tracing

Per-attempt ledger spans are design-stage: the intended shape is one outbound span per attempt named sbproxy.ledger.redeem, carrying sbproxy.ledger.idempotency_key and W3C TraceContext on the outbound request so the trace stitches end-to-end with a ledger that emits OTel spans. The HTTP ledger client does not emit those spans or inject traceparent today.

What ships now: exemplars on sbproxy_ledger_redeem_duration_seconds_bucket carry the active trace id, so Grafana can jump from "this latency outlier" straight to the matching trace in Tempo.

Limitations

  • Detection is User-Agent based by default. Crawlers that lie about their UA bypass the check unless reverse-DNS or Web Bot Auth signals catch them; layer this with bot-detection or WAF policies for defense in depth.
  • The in-memory ledger is single-process. Multi-replica deployments without an HTTP ledger need sticky session affinity to one replica.
  • content_shape is advisory. The field flows through metrics and the redeem payload but is not yet used as a tier filter.
  • Per-agent pricing requires the agent-class resolver to be enabled; the resolver runs unconditionally by default, but operators who explicitly disable it fall back to UA-only matching and lose the per-vendor distinction.

See also

  • configuration.md - schema reference.
  • payment-settlement.md - proxy.payments: rails, durable state, timeouts, reconciliation, and the exact unsupported boundaries.
  • 402-challenge.md - the exact challenge, credential, error, and receipt bytes.
  • ai-gateway.md - how this policy interacts with ai_proxy upstreams.
  • observability.md - metrics, logs, traces, dashboards.
  • examples/ai-crawl-control/ - runnable example.