Awesome AI Tokenomics [](https://awesome.re)
August 3, 2026 · View on GitHub
A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.
Every entry is a link with a one-line summary: what it does, and the number behind it. On top of the list sit a few short pages written here: practices (what to do), concepts (how the economics work), claims (what we currently believe, with the evidence), and setups (configs you can paste straight into Claude Code or Codex). It's a reference to browse, grep, or hand to your agent - not a product.
"I feel nervous when I have subscription left over. That just means I haven't maximized my token throughput."
Andrej Karpathy, No Priors (2026)
As of 2026-07: ~200 entries across five areas.
Topics: Caching · Compression · Context engineering · Memory · Routing · Multi-agent systems · Gateways · Observability · Benchmarks · Cache accounting · Budgets · Pricing models · Energy
Contents
- Where to start
- Legend
- Monitor
- Optimize
- Govern
- Understand
- Measure
- Practices
- Concepts
- Claims
- Setups and skills
- Related lists
Where to start
Just want the numbers: the five area sections below hold every entry. Want the method: read the practices first, then the concepts behind them. Building something: setups and skills holds runnable configurations.
Legend
Each entry ends with a kind badge: (blue, with the license when known), or a gray badge for
,
,
,
for companies, and
. Plain entries are articles. GitHub-hosted tools also carry a live last-commit badge.
Monitor
Dashboards
- ccusage - An open-source CLI that reads local agent logs to report token usage and cost across 15 coding-agent sources, with caching-aware pricing.
- Claude Code Usage Monitor - A live terminal dashboard for Claude Code usage, with burn-rate analytics, P90 limit detection, and session-expiry forecasts.
- claude-usage - A local dashboard for Claude Code token usage, costs, and session history; Pro and Max subscribers get a quota progress bar.
- ClaudeBar - A macOS menu-bar app that monitors AI coding quotas across 11 providers; the README declares MIT but ships no license file, so the OSS grant is unconfirmed.
- CodeBurn - An open-source tracker for 36 coding tools whose optimize command flags named harness-waste patterns with dollar estimates it later checks against actuals.
- Codex Usage Tracker - A local-first dashboard, CLI, and MCP tools indexing Codex CLI logs into SQLite to show where tokens, credits, and cost go, including cache ratios.
- CodexBar - A free, open-source macOS menu-bar app that shows limits and reset timers at a glance across dozens of AI providers, plus credit balances and spending.
- CodeZeno Usage Monitor - A Windows taskbar widget showing real-time Claude Code quota and usage at a glance, without opening a terminal.
- Datadog LLM Observability - Cost - Datadog's LLM Observability estimates per-request cost across 800+ models from token counts and public pricing; invoice reconciliation is a separate product.
- gh-aw (GitHub Agentic Workflows) - GitHub's agentic-workflows runtime with first-party per-run token and cost metering, plus budget caps that stop a workflow mid-run.
- Grafana Cloud GenAI Observability - Grafana Cloud's GenAI Observability ships a prebuilt dashboard for LLM cost, token usage, and latency, built on top of the OpenLIT SDK.
- OpenLIT - An open-source (Apache-2.0), OpenTelemetry-native platform with a self-hosted dashboard for LLM cost, token, and latency observability.
- OpenUsage - A native Swift macOS menu-bar meter for 10 AI coding subscriptions, showing session and weekly limits, credits, and estimated spend from local credentials.
- TokenTracker - A local-first token and cost dashboard for 27 coding tools, with a desktop pet, native widgets, and achievements as a distinct gamified take on usage metering.
eBPF Kernel Capture
- AgentSight - Uses eBPF to watch an AI agent from the kernel boundary, correlating what it said it would do with what it did, with under 3% overhead.
- OpenTelemetry eBPF Instrumentation (OBI) - GenAI / MCP - OBI is OpenTelemetry's zero-code eBPF instrumentation (formerly Grafana Beyla) that captures GenAI and MCP traces at the kernel layer with no SDK.
Observability
- Langfuse - An open-source platform for tracing, evaluating, and analyzing LLM and agent transcripts, with a prompt-management layer on top.
OTel for LLMs
- OpenLLMetry - An open-source set of OpenTelemetry-based SDKs and instrumentations, built by Traceloop, for LLM apps.
- OpenTelemetry GenAI Semantic Conventions - OpenTelemetry's GenAI Semantic Conventions define the vendor-neutral token, cost, and cache attribute names that OpenLLMetry and Phoenix both converge on.
Tracing
- Arize Phoenix - A source-available (Elastic License 2.0) LLM tracing platform recording per-span token counts and USD cost via OpenTelemetry.
- claude-tap - A local trace viewer intercepting API traffic from 14+ coding agents, showing per-request token breakdowns: input, output, cache read, cache creation.
- LangSmith - Cost Tracking - LangSmith is LangChain's commercial LLM/agent observability SaaS.
- Opik - Comet's open-source (Apache-2.0) LLM observability platform, with per-span USD cost estimated from token usage.
Optimize
Caching
- GPTCache - An open-source semantic cache returning a stored LLM response for a paraphrased repeat query via vector search, skipping the paid call.
- khazad - A transport-layer semantic cache for LLM APIs on Redis 8 Vector Sets: it intercepts HTTP traffic with zero application code changes and replays cached responses.
- LMCache - A self-hosted KV-cache layer beneath vLLM, giving token-level cache-hit observability for teams who own their GPUs, not a hosted bill.
- prompt-cache - A Go LLM proxy that adds a three-tier semantic cache: high similarity hits directly, low skips, and a gray zone runs a cheap verification model.
- Redis LangCache - Redis's fully managed semantic cache: a REST API that returns a stored response when a new query is similar to a past one.
Cheap Local Models
- llama.cpp - The foundational open-source (MIT) local LLM inference engine most of the local ecosystem runs on, with an OpenAI-compatible server built in.
- Ollama - A runtime for running open-weight models like Qwen, DeepSeek, and GLM-5.1 locally, shifting inference onto hardware you already own.
Compression
- Context Mode - This MCP server sandboxes tool calls and returns only the distilled result, claiming a 98% cut: 315 KB of output down to 5.4 KB.
- headroom - An Apache-2.0 context-compression tool for LLM/agent pipelines, the category's largest repo, confirmed organic by star-forensics.
- LLMLingua - Microsoft's prompt-compression library that uses a small model to drop low-information tokens before a prompt reaches the target LLM.
- llmtrim - A local proxy that compresses a coding agent's prompt, tool schemas, and history before forwarding and can reroute Claude calls to Grok.
- Minification of state-in-context agents - the clean waste-vs-capability datapoint - This ICPC 2026 study found that minifying code in a coding agent's context cuts input tokens by 42% but costs 12 percentage points of accuracy.
- rtk - A single-binary Rust CLI proxy that intercepts and compresses the output of common dev commands before it reaches an LLM coding agent's context window. Its headline figures are token reduction, not measured cost reduction.
- TOON (Token-Oriented Object Notation) - TOON is a compact, human-readable, lossless serialization of the JSON data model, designed for LLM input.
Context Engineering
- AgentDiet - trajectory reduction ("Reducing Cost of LLM Agents with Trajectory Reduction") - AgentDiet is an inference-time module that strips useless, redundant, and expired information from an agent's trajectory, without hurting performance.
- Agentic Context Management for long-horizon tasks - ACM reduces peak token usage by around 14-20% while raising accuracy, but trajectories get longer: in one benchmark peak fell 14% while tool calls went up 2.4x, and the paper reports no total-token or dollar result - window pressure is not the bill.
- Anthropic vendor-native context management (context editing + memory tool + server-side compaction) - Anthropic's context editing, memory tool, and server-side compaction cut token consumption by a vendor-reported 84% in a 100-turn web-search evaluation.
- Claude Code compaction engine - the three-tier mechanism and its cache/correctness failure modes - Claude Code's harness - not the API - decides how to trim a filling context window, and it does this through a three-tier compaction engine.
- Codex CLI compaction cost - over-eager compaction as a token-amplification loop - Upgrading Codex CLI from v0.116 to v0.118 made context compaction fire twice as often, doubling or tripling token consumption for identical tasks.
- Context Rot - LLM performance degrades as input length grows - This is Chroma's controlled study of how LLM output quality changes as input length grows, holding task difficulty fixed.
- ContextBudget - context management as a budget-constrained sequential decision - ContextBudget's BACM method has an agent decide when and how much to compress its history based on remaining context budget, not a fixed rule.
- Cursor vendor-native context management - dynamic context discovery + Composer self-summarization - Cursor's dynamic context discovery loads tool schemas and large outputs on demand instead of eagerly, a change the vendor reports cut context usage by 46.9%.
- Repomix - Packs an entire repository into a single AI-friendly file, reporting token counts and using Tree-sitter to compress code to signatures only.
- RULER - NVIDIA's RULER benchmark found that of models claiming 32K+ token context windows, only half actually maintain quality once you fill them to 32K.
- Self-Compacting Language Model Agents - This paper introduces SELFCOMPACT: instead of fixed-interval summarization, the model itself decides when and how to compress a growing agent trace.
- Serena - An open-source (MIT) MCP toolkit that gives a coding agent IDE-grade semantic code retrieval and editing: 'the IDE for your coding agent'.
Cost Controls
- Claude Code spend-governance bundle (v2.1.216-219) - caps tightened, fan-out default loosened - Claude Code's v2.1.216-219 bundle (2026-07-20 to 07-24) hardens spend controls - a concurrent-subagent cap and enforced --max-budget-usd.
- Harness-side runaway-loop cost guardrails (Claude Code + Codex, July 2026) - In mid-July 2026 Claude Code and Codex both shipped first-party guardrails against runaway agent loops within days of each other.
Gateways and Proxies
- Bifrost (Maxim AI) - Bifrost is a Go-based AI gateway fronting 1,000+ models that measured just 11 microseconds of added latency per request at 5,000 requests per second.
- Cloudflare AI Gateway (Spend Limits) - Cloudflare AI Gateway is an edge-native LLM proxy that added dollar-denominated spend limits in June 2026, blocking or rerouting requests once a budget is hit.
- Helicone - An open-source (Apache-2.0) LLM proxy that logs every request's cost, latency, and tokens in one line of code; Mintlify acquired it in March 2026.
- Kong AI Gateway - The AI layer of Kong's API-gateway platform: a proxy that meters LLM/agent/MCP traffic for billing, showback, and chargeback.
- LiteLLM - An open-source gateway fronting 100+ LLM APIs that computes real per-request dollar cost from a live pricing map, with spend limits.
- OpenRouter - A unified API gateway fronting 400+ models across 70+ providers that auto-routes each request by price, with fallback on outages.
- Portkey AI Gateway - Routes LLM traffic across providers and enforces hard USD budget limits on virtual keys, auto-expiring a key once its cap is hit.
Memory
- claude-code-memory-setup - A practitioner recipe pairing an Obsidian memory vault with a local AST code-graph tool.
- claude-mem - A coding-agent observational-memory layer that captures every session, compresses it with AI, and re-injects relevant context next time.
- Cognee - An open-source (Apache-2.0) AI-memory platform giving agents persistent memory via a self-hosted knowledge graph, via remember/recall/forget.
- Karpathy's LLM Wiki - Andrej Karpathy's LLM Wiki pattern has an agent build and maintain a persistent markdown wiki from your sources, instead of re-retrieving raw files.
- LangMem - LangChain's long-term memory library: it extracts and consolidates facts from conversations and integrates natively with LangGraph's memory store.
- Letta (MemGPT) - The MemGPT lineage project: a platform for stateful agents that pages an LLM's context like an OS.
- Mem0 - An open-source memory layer that extracts salient facts from conversations and retrieves only the relevant ones per call, not the full history.
- Supermemory - A memory and context engine that self-reports 95% recall on LongMemEval while adding only ~720 tokens of context.
- Zep / Graphiti - A memory platform for agents built on temporal knowledge graphs; it self-reports serving benchmark answers from a few thousand tokens of retrieved context. (also: Graphiti (OSS engine) · zep repo)
Multi-Agent Systems
- Framework orchestration overhead (the manager-LLM tax) - Hierarchical frameworks like CrewAI add a manager-LLM delegation tax, an extra model that plans and validates, though its cost is unquantified in any primary.
- SupervisorAgent - "Stop Wasting Your Tokens" (runtime supervision) - SupervisorAgent is a lightweight, modular framework for runtime, adaptive supervision of multi-agent systems.
Prompt Agent Loop
- LOOP Skill Engine - LOOP records an agent's first run of a repetitive task with full LLM reasoning, then replays the extracted tool-call template without calling the LLM again.
- Orchestrator-worker model tiering (frontier plans / cheap executes) - A capable model plans while cheaper agents execute; the pattern now ships as a vendor default, hitting 89.7% of LLM quality at 4% of the cost.
- token-ninja - Intercepts deterministic commands like git status or npm test before they reach the model, running them locally and skipping the LLM call.
Retry and Reliability
- Claude Code v2.1.199 - transient-retry hardening & partial-output preservation - Claude Code v2.1.199 now auto-retries rate-limit errors with backoff and raised the default retry ceiling to 300, up from a prior cap of 15.
Routing Model Selection
- Antigravity CLI - per-subagent model-tier routing + /effort (v1.1.5) - Antigravity CLI v1.1.5 shipped first-party per-subagent model-tier routing (a model: flash|pro field in custom-agent frontmatter) plus an /effort control.
- Claude Code Router - A local gateway that puts Claude Code, Codex, and other coding CLIs behind one endpoint and routes each request by ordered condition rules, Node.js script rules, or a prompt tag that lets the agent pick a model per subagent. The project publishes no savings figure, and routing scripts run as fully trusted code next to your credentials - only use scripts you wrote yourself.
- Claude Code via a LiteLLM gateway (cheap-tier-in-front setup) - Pointing Claude Code's ANTHROPIC_BASE_URL at a local LiteLLM proxy lets cheaper or non-Anthropic models absorb work the frontier model would otherwise bill for.
- Cluster, Route, Escalate - cost-aware cascaded serving - This paper proposes a two-stage cost-aware cascade for LLM serving that combines routing and escalation into one framework.
- Cursor Router - Cursor's Auto mode classifies each request and routes it to a model under three modes (Intelligence, Balance, Cost), with reported savings measured cache-miss-inclusive. All percentages are Cursor's own, against a constructed all-Opus baseline, with no third-party replication yet.
- Distilling agent behavior into small task-specific models - Distilling a large agent's behavior into a small 0.5-3B model lets most of its work run at a fraction of the frontier model's per-token cost.
- GitHub Copilot auto model selection - Copilot's Auto setting routes by real-time model health and task complexity, and only along cache boundaries: GitHub states mid-session model switching "has shown increased cost without ample improvements in quality." The 10% discount for paid plans in Auto is a pricing multiplier, not a measured routing saving.
- MTRouter - per-turn cost-aware routing with history-model joint embeddings - MTRouter picks a different model for each turn of a multi-turn conversation, rather than one model per query, to hit a cost budget without losing quality.
- Not Diamond - Not Diamond's meta-model predicts, per input, which LLM will give the best answer at the lowest cost, then routes the request there.
- OpenCode - explicit cost-tier routing - OpenCode is an open-source (MIT) coding-agent CLI with its own explicit cost- and model-routing configuration, set directly in config.
- opencode-fusion - An OpenCode config layer that denies the main agent's edit and search tools so they are removed from its tool schema entirely, forcing every file change through a cheaper sidekick agent. Model assignments are fixed per role at startup, and the project publishes no savings measurement of its own.
- OrcaRouter - production LinUCB bandit router (hybrid offline-online) - OrcaRouter is a production LLM router built on a LinUCB bandit, with its cost/quality tradeoff independently confirmed on the RouterArena leaderboard.
- Plano (formerly archgw) - An Envoy-based proxy whose router matches queries to user-defined domains and actions via a small routing model, rather than picking by benchmark rank. Since July 2026 it also prices the warm cache a model switch would discard, and vetoes switches once their cumulative cost passes a configured overhead cap.
- RouteLLM - LMSYS's open-source router sending each query to a cheap or expensive model based on a trained cost threshold, as a drop-in server.
- ruflo (formerly Claude-Flow) - cost-adjusted model routing - ruflo is an open-source agent meta-harness for Claude Code and Codex, providing swarm orchestration and persistent memory. Ships on npm as
claude-flow(v3.17.0). - vLLM Semantic Router - Sends routine queries to cheap or local models and hard ones to stronger backends, as an open-source, self-hostable router.
- Weave Router - A drop-in proxy that picks a model for every request with an on-box embedding cluster scorer derived from Avengers-Pro, speaking all three provider APIs (BYOK, OTLP traces, one-command setup for Claude Code, Codex, and opencode). Source-available under Elastic License 2.0, which bars offering it as a hosted service; its cost-reduction figures are vendor-reported, not independently measured.
Search and Retrieval Boundary
- Exa - A search API for agents that bills content retrieval separately per type, so an agent can buy query-scoped highlights or a summary instead of full page text.
- Firecrawl - A web scraping API that converts pages to markdown or structured JSON before they reach the model, billed at 1 credit per page.
- Parallel (parallel.ai) - A web search and extraction API for agents that publishes an accuracy-versus-cost table across five benchmarks with sample sizes, judge model and test dates stated.
- Tavily - A search API for agents that returns capped content snippets instead of pages, priced at $0.008 per credit with one credit per basic search; acquired by Nebius (announced 2026-02-10).
- Valyu - A search and deep-research API whose cost-versus-accuracy results on the third-party DRACO benchmark ship with an open harness, raw outputs and per-provider runners; the harness repo carries no license file.
Serving Inference
- RLM-Cascade - response-level speculative decoding at the gateway - RLM-Cascade, from a PayPal team, has a cheap draft model answer first and an Opus 4.8 verifier accept or rewrite it, at roughly 2% of Opus's cost.
- SGLang - A high-performance serving framework for large language and multimodal models.
- vLLM - The canonical open-source LLM serving engine, using PagedAttention to manage KV-cache memory in blocks so more requests batch at lower cost.
Test-Time Compute
- Stop When Reasoning Converges - Reasoning models often keep generating steps after a solution has already stabilized, wasting tokens and adding latency - a pattern the paper describes as overthinking.
- When more reasoning hurts - the test-time-compute ceiling - Two 2026 papers found giving a model more reasoning budget makes it perform worse and cost more; past a point, tool delegation wins outright.
Tool Protocol Overhead
- Code execution with MCP (Anthropic) - Anthropic proposes agents call MCP servers by writing and executing code instead of a tool call per step, so unused tool schemas skip the context window.
- Coral - Gives agents one SQL interface over APIs and internal systems instead of many MCP servers; its own 82-task benchmark reports 64% fewer tokens on the complex-task slice, 41% across all tasks.
- MCP Tool Descriptions Are Smelly! - This study found poorly-written MCP tool descriptions measurably hurt agent efficiency, using an LLM-jury scanner and an A/B protocol on MCP-Universe.
- StackOne Falcon - An execution engine that cuts tool-calling tokens by filtering tool definitions, shaping responses and running code at the edge, with search-first discovery benchmarked on 1,843 tasks; its broader reduction percentages are vendor-reported.
- Tool Attention Is All You Need - MCP re-sends every tool's full schema on every turn, whether or not the agent needs it - a protocol tax known as the MCP/Tools Tax.
Govern
Allocation Chargeback
- CloudZero - An established commercial cloud and AI cost-intelligence / FinOps platform that brands itself 'The AI ROI Company'.
- JetBrains AI moves business plans from monthly licenses to 12-month credits - JetBrains is moving business AI from monthly per-seat licenses to 12-month reallocatable credits plus a governance dashboard.
- Mavvrik (fmr. DigitalEx) - Mavvrik is an AI/hybrid-infrastructure cost governance and FinOps platform, rebranded from DigitalEx in February 2025.
- Pay-i - An SDK-based GenAI cost-observability platform that tracks token-level spend per call and rolls it up into cost-center allocation across orgs and apps.
Anomaly Detection
- Denial-of-Wallet / token-exhaustion attacks - Denial-of-wallet attacks exploit pay-per-token pricing to inflate a bill, via stolen-credential LLMjacking or agents steered into runaway token use.
- Governance Decay - compaction silently erasing safety/governance constraints - Compacting an agent's context can silently erase governance rules: across 7 model families, violations rose from 0% to 30%, up to 59% for some.
Billing Audit FinOps
- FinOps for AI - canonical practitioner framework for governing AI/LLM spend - FinOps for AI is the FinOps Foundation's official practitioner framework for governing AI, GPU, and token spend.
- Vaudit - TokenAudit - Vaudit is an AI-native, independent spend-auditing and recovery platform (San Francisco, founded late 2023). TokenAudit is its LLM invoice-reconciliation product.
Budgets Caps
- Claude Code's 5-hour/weekly usage quotas - Anthropic has stopped publishing exact numbers - Anthropic stopped publishing exact Claude Code usage quotas, describing Max plans only as 5x/20x multipliers of Pro with no absolute numbers.
- TrueFoundry (AI Gateway - Budget Limiting) - TrueFoundry is an enterprise GenAI deployment/gateway company founded by ex-Meta founders.
Energy Carbon
- CodeCarbon - An open-source (MIT) library for estimating a workload's energy use and CO2e emissions, and ML's widely-cited carbon baseline.
- EcoLogits - Estimates the energy and carbon footprint of calling generative-AI APIs: the hosted counterpart to CodeCarbon, which measures your own hardware.
- Epoch AI - how much energy a query uses (the per-token energy anchor) - Epoch AI built a transparent, first-principles estimate of how much energy one LLM query costs.
- Google - measuring the environmental impact of AI inference (provider disclosure) - Google published a first-party disclosure of the energy, carbon, and water cost of a median Gemini Apps text prompt, authored by Amin Vahdat and Jeff Dean.
- ML.ENERGY Leaderboard - Version 3.0 of this leaderboard measures real GPU inference energy across 46 models x 7 tasks, finding reasoning models use roughly 25x the energy of others.
Policy Enforcement
- ActPlane - An eBPF-based, OS-level policy-enforcement engine for AI-agent harnesses like Claude Code and Codex.
- AEGIS - An open-source (MIT) pre-execution firewall and cryptographic audit layer for AI agents.
- GitHub Copilot default model enablement - From 2026-08-26, new GA models are enabled by default for Copilot Business/Enterprise orgs under a single opt-out policy, inverting the prior per-model opt-in; open-weight models and models outside GitHub's data-retention agreement (e.g., DeepSeek, Kimi K2.7, Fable 5) stay excluded regardless of the policy.
- MCPGuard-Dynamic - An early-stage, research-grade kernel-level eBPF sandbox for MCP (64★), published under Meta's official GitHub org.
Spend Management
- ChatGPT Enterprise - usage analytics & spend controls - OpenAI's first-party spend layer for ChatGPT Enterprise/Business: a Global Admin Console with credit caps, request workflows, and a Cost API.
- Claude Enterprise - admin analytics & cost controls - Anthropic's first-party spend surface for Claude Enterprise/Team admins: org-level spend caps, model defaults, and per-user cost analytics via the Admin API.
- PointFive (AI Efficiency OS / TokenShift) - PointFive's TokenShift governs coding-agent token spend across Claude Code, Cursor, Codex, and more, claiming a 10-20% cut across 11 partners.
- Revenium - Tracks AI agent spend at runtime to the cent, attributing every model call and tool cost to its workflow, with auto-shutoff on runaway budgets.
- Vantage - A FinOps platform ingesting native token-level cost data from Anthropic and OpenAI's own usage APIs, plus Cursor and cloud spend.
- Vercel AI Gateway - per-API-key budgets - Vercel AI Gateway lets you cap spend per API key in dollars (min $1) with a daily/weekly/monthly refresh, rejecting further requests once the cap is hit.
Understand
Buyer Incentives
- Claude "subscription arbitrage" and its (announced, then paused) end - Users route agentic workloads worth far more than a subscription's price through cheap Pro/Max plans; Anthropic tried to close this, then paused the fix.
- Coding-agent native spend controls (2026) - Within six weeks in 2026, Cursor, GitHub Copilot, and OpenAI each shipped native admin spend controls: budget caps, credit metering, usage dashboards. (also: GitHub Copilot · OpenAI)
- GitHub Copilot metered-billing bill-shock - the demand-side reaction ("tokenpocalypse") - GitHub's move from flat-rate Copilot plans to metered AI Credits exposed agentic workflows' true per-token cost and triggered a mass bill-shock backlash.
- Lanai - Lanai's AI @ Work platform discovers every sanctioned and shadow AI workflow across an org and maps its token spend to the KPIs it actually drives.
- State of FinOps 2026 - AI spend management is now the norm - This is the FinOps Foundation's sixth annual State of FinOps survey, the practitioner census of how organizations manage cloud and AI spend.
- The "$47k Claude Code bill" - the anchor bill-shock anecdote and its mechanistic debunk - A viral $47,000-in-90-days Claude Code bill story was debunked by a teardown pinning the real driver on quadratic context re-ingestion, not runaway use.
- Uber caps AI-coding spend at $1,500/mo per tool after burning its budget in ~4 months - Uber capped AI-coding spend at $1,500 per employee per tool after burning its entire annual budget in roughly four months of encouraged maximal use.
Compression Efficacy
- JetBrains independently measured two token-saving skills against their own claims - JetBrains independently A/B-tested two token-saving skills: rtk ran +7.6% more expensive at low effort (claimed 60-90% cut), Caveman saved ~8.5% (claimed 65%). (also: Caveman A/B post)
Consolidation
- Cisco acquires Galileo (LLM eval/observability) → folded into Splunk Observability - Cisco acquired Galileo, an LLM/agent evaluation and observability platform, folding it into Splunk Observability Cloud's AI Agent Monitoring.
- Tokenomics Foundation (Linux Foundation + FinOps Foundation) - The Linux Foundation is launching the Tokenomics Foundation to build open standards for AI token spend, extending FOCUS to cover token-based costs.
Market Competitors
- Aider - an OSS coding CLI that meters its own dollar cost - An open-source terminal coding agent (Apache-2.0, ~48k★) with built-in per-message dollar-cost tracking and a public polyglot leaderboard.
- Amp (Sourcegraph) - pay-as-you-go, no-markup pricing + mode-based routing - Amp, Sourcegraph's coding agent, passes through LLM cost with zero markup for individuals and teams, with a cost/capability mode: Deep, Smart, or Rush.
- Cline - an OSS coding agent on a bring-your-own-key cost model - An open-source AI coding agent (Apache-2.0, ~65k★, ~4.8M VS Code installs), built on a bring-your-own-API-key cost model.
- Factory (droids) - subscription pricing + $150M Series C at $1.5B - Factory prices its Droids agents as flat subscription tiers ($20/$100/$200/mo) with usage-based rate limits, not per-token metering, after a $150M round.
- Gemini CLI retirement → Antigravity CLI (open-source coding agent closes, pricing restructures) - Google retired the open-source Gemini CLI on 2026-06-18, pushing users onto closed-source Antigravity CLI and $100/$200-per-month paid tiers.
- OpenHands - MIT OSS coding agent, free local + free cloud tier, at-cost LLM option - An MIT-licensed open-source coding-agent platform with a free local mode, a free cloud tier, and an at-cost LLM pricing option.
Market Sizing
- AI Tokenomics: how to tokenmin while ROImaxxing (MMC Ventures) - MMC's 2026-06-30 map of the token-efficiency vendor landscape across five levers: context and memory, multi-model systems, inference optimisation, routers and gateways, and output optimisation. Its headline waste estimates are the firm's own, without published methodology.
- Gartner - worldwide AI spending forecast: $2.59T in 2026 (+47% YoY) - Gartner's latest forecast puts worldwide AI spending at $2.59 trillion in 2026, up 47% year-over-year, with infrastructure over 45% of the total.
- Menlo Ventures - enterprise generative-AI spend $11.5B → $37B (2024→2025) - Menlo Ventures found enterprise generative-AI spend hit $37B in 2025, up 3.2x from 2024, with coding tools the largest application category at $7.3B.
Model Economics
- "Qwen 3.6 27B is the sweet spot for local development" - Migdał / Quesma (first-party) - Piotr Migdał's #1-on-Hacker-News essay argues Qwen3.6-27B (dense) is the first local model good enough for real coding instead of a metered cloud API.
- Artificial Analysis - Coding Agent Index - This benchmark scores full model-plus-harness stacks, spanning $0.27 to $11.80 per task: a roughly 44x range at similar quality, per Artificial Analysis.
- Artificial Analysis - Intelligence Index + Blended Price - Artificial Analysis's Intelligence Index plots a 0-100 capability score against blended price per million tokens, live across 85-122 base LLMs.
- Claude Opus 5 - flat price vs Opus 4.8, but 1M context and thinking on by default - Claude Opus 5 launched 2026-07-24 at the same $5/$25 per MTok as Opus 4.8, but ships 1M context and thinking on by default.
- Gemini 3.6 Flash - a Flash tier marketed on fewer tokens per task, not just a lower unit price - Gemini 3.6 Flash (2026-07-21) is priced lower than 3.5 Flash at $1.50/$7.50 per MTok and uses 17% fewer output tokens on the same work.
- Kimi K2.6/K2.7-Code and GLM-5.2 official API pricing - Kimi K2.7-Code ($0.95/$4.00 per million tokens) and GLM-5.2 ($1.40/$4.40) undercut Claude Sonnet 5 and GPT-5.5 on raw price by 2-5x; Moonshot's flagship is Kimi K3 at $3/$15, and GPT-5.6 Luna holds the bottom of the hosted price curve since its 2026-07-30 cut.
- Local / open-model economics for coding - state of the field (2026) - Open-weight coding models now score 77-81% on SWE-bench Verified, within a few points of closed frontier models, reshaping self-host-vs-API math.
- OckBench - measuring token efficiency / verbosity of LLM reasoning - OckBench answers a specific tokenomics question: which model burns the most tokens for the same answer?
- Reasoning-token billing across providers - the hidden output multiplier - Every major AI provider bills a model's hidden reasoning tokens at the most expensive output rate, without ever returning them to the caller. (also: Google · DeepSeek)
- Reasoning-token consumption behavior - length ≠ effort, and verbosity is a separate lever - Chain-of-thought can burn about 258 tokens on problems a direct answer solves in 15 (roughly 17x overhead), and simple agentic steps trigger it by accident.
Pricing Models
- Batch / Priority / Flex service tiers - the scheduling axis of token pricing (clustered, cross-vendor) - Every major LLM vendor sells the same lever, trading latency for price via async batch scheduling, with Anthropic, OpenAI, and Google all near 50% off. (also: OpenAI · Google)
- Bessemer - the AI pricing & monetization playbook (seat → usage → outcome) - Bessemer's playbook argues AI pricing is shifting from per-seat to consumption/outcome-based, citing Intercom's $0.99-per-resolved-ticket model.
- Cached-input discounts - the ~90%-off lever behind cache-accounting - Cache-read pricing discounts input tokens by about 90%: the biggest lever on an agentic bill, since input is roughly 85% of session cost.
- ChatGPT subscription tiers and Codex CLI bundling/pricing (2026) - OpenAI bundles Codex CLI into every ChatGPT tier from Free through the new $100/mo Pro plan, differing only by rate-limit multiplier.
- ChatGPT workspace-agent credit billing (effective July 6, 2026) - OpenAI ended the free preview for agent runs invoked inside ChatGPT Business, Enterprise, Edu, and Teachers on 2026-07-06.
- Cursor charges by tokens, split into first-party and third-party pools - Cursor meters by tokens per million (input/output/cache-write/cache-read), split into a first-party pool and a third-party API pool. (also: Teams pricing blog)
- Devin's Agent Compute Unit has no published definition of what it meters - Devin bills Enterprise usage in Agent Compute Units, but no official doc defines what an ACU measures (not tokens, seconds, or calls).
- Doubleword - Sells async and batch inference on open models, publishing a cost-per-1B-tokens table that holds capability constant using the third-party Artificial Analysis index.
- Fable 5 leaves subscription inclusion - frontier tier moves to usage-credit metering (July 7 cliff) - Fable 5's subscription saga settled 2026-07-20 (after two extensions) as a primary-confirmed two-tier split.
- Google AI Pro price and Gemini/Antigravity free-tier limits (2026) - Google AI Pro is confirmed at $19.99/month, beneath the $99.99 and $199.99 AI Ultra tiers giving higher rate limits on the Gemini API and Antigravity.
- GPT-5.6 family (Sol / Terra / Luna) - API pricing - OpenAI's GPT-5.6 family prices three tiers: Sol at $5/$30 per million tokens, Terra at $2/$12, and Luna at $0.20/$1.20 (Terra and Luna cut 2026-07-30, three weeks after launch). "Fast mode" is the renamed priority tier at a flat 2x.
- LiteLLM flex/priority service-tier cost keys - the harness-level tier-routing lever - LiteLLM automatically prices requests made at a non-standard tier like flex or priority, applying the right discounted or premium rate automatically.
- LLM price decline + Jevons paradox - unit price crashes, total spend climbs - Per-token prices are falling roughly an order of magnitude per year, while total AI spend rises even faster.
- LLM token pricing dimensions - the structure of a token bill - This maps out how frontier LLM APIs meter and price tokens, read straight off the two largest providers' pricing pages, Anthropic and OpenAI.
- OpenAI is winding down the self-serve fine-tuning API and platform - OpenAI is winding down self-serve fine-tuning because prompting got cheaper and more capable than fine-tuning for most uses, cutting off customers by 2027.
- Tokenization multiplicity & overcharging - the pay-per-token integrity problem - Two academic papers show the same output can be billed a different token count depending on tokenization, and providers can be incentivized to inflate it.
- Windsurf became Devin Desktop and switched credits to token-based quota - Windsurf became Devin Desktop and in March 2026 swapped opaque per-model credit multipliers for token-based quota where free models cost nothing.
Reliability SLAs
- Reserved-capacity reliability economics (Azure PTU · AWS Bedrock MU) - Azure's Provisioned Throughput Units and AWS Bedrock's Model Units both let buyers reserve guaranteed capacity, billed hourly whether or not it's used.
Unit Economics
- Cloud Capital - gross margin in the age of AI (the vendor/supply side) - AI-native software runs at roughly 50-60% gross margin versus 70-80% for SaaS, since inference and compute became a large, variable cost of goods sold.
- Cost-of-Pass - an economic framework for evaluating language models - Cost-of-Pass defines the expected dollar cost of one correct answer as inference cost divided by success rate, pricing benchmark accuracy directly.
- DORA 2025 - AI as amplifier, and the delivery-stability tension - Google's DORA program found that as AI adoption becomes universal, delivery throughput rises but so does instability: AI as an amplifier, not a pure win.
- Faros - "The Acceleration Whiplash" (AI Engineering Report 2026) - The "velocity has a hidden bill" study: telemetry from 22,000 developers across 4,000 teams over two years.
- getDX - AI coding assistant pricing & ROI guide (2026) - Typical AI coding tools cost $200-600 per engineer monthly in seat plus token spend, per getDX, for a median 7.76% PR gain: below vendors' claimed 3-10x.
- METR - measured vs perceived AI productivity (the RCT) - METR's controlled trial found developers took 19% longer to finish issues when allowed to use AI, while still believing it had sped them up by 20%.
- Paid (paid.ai) - A monetization platform for AI agents that sets pricing, tracks delivery cost per action and reports margin per customer; distinct from the similarly named Pay-i.
Measure
Benchmarks Evals
- Claw-SWE-Bench - Found that adapter/harness design alone swings an agent's Pass@1 score by about 54 percentage points on the identical model backbone.
- Coding Benchmarks Are Misaligned with Agentic SE (Tessl) - A position paper from Tessl (London, UK) argues that today's coding benchmarks don't measure what people think they measure.
- Deterministic Anchoring - how much static structure do code agents need? - This ISSTA 2026 paper found injecting static-analysis facts as plain-text comments raises a code agent's Pass@1 by 3.4pp and cuts trajectories by 1.6 rounds.
- GitHub Copilot agentic-harness efficiency evaluation (first-party offline ablation) - GitHub's own benchmark plots Copilot's agentic-harness resolution rate against dollar-cost-per-task across five benchmarks and four frontier models.
- Harness-Bench - Holds the task, model, and budget fixed while varying only the agent harness, across 5,194 trajectories spanning 6 harnesses and 8 models.
- LoCoMo - Snap Research's very-long-term conversational-memory benchmark; the canonical dataset that Mem0, Zep, and Supermemory all cite, kept as a background instrument.
- LongMemEval - The peer-reviewed (ICLR 2025) benchmark for long-term memory in chat assistants.
- MemoryBench - A pluggable harness to run memory systems (Supermemory, Mem0, Zep) head-to-head across datasets like LoCoMo; useful for standardizing comparison.
- Prompt Compression in the Wild - the end-to-end referee for compression - This ECIR 2026 study found LLMLingua's compression yields up to 18% speed-up only in a narrow window; outside it, the compression step cancels the gains.
- promptfoo - An open-source CLI/CI harness for testing LLM prompts and agents that records per-eval token usage and cost as an assertable metric.
- RedundancyBench - can anyone even detect a redundant step? - RedundancyBench is a benchmark for step-level redundancy detection in agent trajectories - can a model even spot the wasted step in an agent's history?
- RouterArena - An open evaluation platform and live leaderboard for LLM routers - systems that auto-select a model per query.
- SWE-bench - The canonical accuracy-only software-engineering benchmark (2,294 real GitHub issue tasks, 12 Python repos, ICLR 2024). The cost-aware derivatives in this list build on it.
- SWE-Effi - cost-aware re-ranking of SWE-agents under resource budgets - SWE-Effi re-ranks popular AI issue-resolution systems on a SWE-bench subset by cost-under-resource-constraints instead of by accuracy alone.
- Terminal-Bench - The canonical benchmark for AI agents in real terminal/CLI environments, with 89 tasks each vetted through ~3 reviewer-hours.
Cache Accounting
- Gemini context caching - the per-hour storage meter (the third-vendor axis) - Among major providers, only Gemini bills a per-hour storage meter for explicit prompt caching.
- OpenAI prompt-caching
cached_tokensaccounting - This is the measurement substrate for prompt-cache savings on the OpenAI API. - OpenAI Responses conversation state - you pay for the whole chain every turn (and the compaction levers) - This entry covers the billing semantics of stateful conversations on OpenAI's Responses API, plus the two server-side compaction levers that mitigate them.
Cost Anatomy
- Anthropic bill anatomy - the whole-bill line-item taxonomy - This is the canonical enumeration of every meter on a Claude API bill, read live from Anthropic's pricing page.
- TensorZero - cross-vendor token-count divergence ("stop comparing $/M tokens") - The same input can produce 2.65x more tokens on one tokenizer than another's: Claude Opus 4-7 runs 1.57x-2.65x more tokens than GPT-5.4 on the same content.
- Vision-token pricing formulas across the big three - Anthropic, OpenAI, and Google each convert an image into billed tokens with a different formula, so no single cross-vendor image-cost number exists. (also: OpenAI · Google)
Harness Overhead
- Agent system-prompt leak corpora (x1xhlol · asgeirtj) - Two maintained collections of extracted agent system prompts and tool schemas across 35+ vendors; the asgeirtj corpus adds harnesses others miss, like Antigravity CLI, under CC0. Raw text with no token counts: a source corpus for overhead measurement, not a measurement.
- Claude Code system prompts (Piebald extraction) - Piebald AI extracts Claude Code's full compiled prompt payload per release: 515 prompt strings and 27 tool descriptions at v2.1.212, each priced in tokens.
- Claude Code vs OpenCode token overhead (Systima study) - Systima measured harness scaffolding overhead before a prompt is even read: Claude Code carries about 32,800 tokens versus OpenCode's 6,900, a 4.7x gap.
Metering
- Cross-vendor coding-agent usage trackers (AgentsView · caut) - AgentsView and caut are open-source tools that read local session logs to aggregate token usage and cost across roughly 20 coding-agent vendors.
- How Do AI Agents Spend Your Money? - This Stanford study is the first systematic look at token spend in agentic coding, running 8 frontier models on 500 SWE-bench Verified tasks.
- OpenCost - AI inference cost tracking - The CNCF Kubernetes cost tool's 2026-07 feature turns GPU and shared-infrastructure spend on vLLM/llm-d deployments into cost per million tokens, reconciled to the infrastructure bill.
- tokview - A local, zero-config proxy showing a coding agent's token spend by session, model, and tool call, flagging re-sent results that multiply the bill.
Transcript Analysis
- Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems (Failure-Aware Observability) - This live waste-detection framework for multi-agent systems found that 58.1% of tokens in failed runs are spent after its first warning fires.
- Faros AI - "Token Intelligence" / Token Attribution Ledger - Faros AI is an enterprise engineering-intelligence SaaS (DORA/SPACE-style productivity analytics).
- token-optimizer (alexgreensh) - token-optimizer is a coding-agent plugin that runs eleven heuristic waste detectors per session and prices the flagged tokens in dollars. Coverage is semi-cross-vendor.
Whole Bill Accounting
- FOCUS 1.4 - the cross-vendor billing-data normalization standard (now with invoice reconciliation) - FOCUS 1.4, the Linux Foundation's billing-data schema, added Invoice Detail and Billing Period datasets to reconcile spend against real invoices.
Practices
Tool-agnostic, evidence-grounded standards for token-efficient agentic coding. Each is one page: TL;DR, claim, evidence, links. Browse the practices.
Concepts
Short reference notes explaining the ideas behind the practices: cache economics, the harness-waste taxonomy, orchestration economics. Browse the concepts.
Claims
Confidence-scored beliefs, clearly labeled as beliefs rather than facts, each with its strongest evidence linked. Read the claims.
Setups and skills
Runnable, validated Claude Code and Codex configurations and skills for token-efficient agentic coding, each labeled with how it was validated. Browse the setups.
Related lists
- Awesome-LLM - Curated papers, frameworks, and resources on large language models.
- Awesome-LLMOps - Tools and platforms for operating LLMs in production.
- awesome-mcp-servers - A collection of Model Context Protocol servers.
- awesome-production-machine-learning - Open-source libraries to deploy, monitor, version, and scale machine learning.
Text content: CC BY 4.0 (LICENSE) · Code and configs: MIT (LICENSE-CODE). Maintained by the team at Quesma.