Built-in pricing catalog

September 14, 2026 · View on GitHub

Baseline verified against official public pricing pages on 2026-08-26 UTC; DeepSeek V4.1 Flash updated for the 2026-09-10 04:00 UTC transition. Prices are USD per one million tokens.

The catalog estimates standard, synchronous, first-party API usage. Batch/flex/priority modes, regional uplifts, negotiated discounts, subscriptions, tool-call fees, taxes, and cache-storage token-hours are excluded unless a row note explicitly says otherwise.

Catalog entries: 132. Generated route rules include provider aliases, context/time tiers, and model-name fallbacks for custom providers.

Family / tierProvider routesModelsInputCache readCache writeOutputMatchNotes
GPT-6 Astra · long contextopenai, openai-codexgpt-6-astra$20$2$25$75prompt ≥ 272,001Standard synchronous tier; long-context rates apply to the full request above 272K input tokens. See https://developers.openai.com/api/docs/models/gpt-6-astra.
GPT-6 Astraopenai, openai-codexgpt-6-astra$10$1$12.5$50prompt ≤ 272,000Standard synchronous tier; long-context rates apply to the full request above 272K input tokens. See https://developers.openai.com/api/docs/models/gpt-6-astra.
GPT-5.6 Sol · long contextopenai, openai-codexgpt-5.6-sol$8$0.8$10$30prompt ≥ 272,000Standard synchronous list-price estimate, including historical calls. Promotional through at least 2026-11-21; no confirmed start or expiry is published. The catalog verification date is not a billing boundary.
GPT-5.6 Solopenai, openai-codexgpt-5.6-sol$4$0.4$5$20prompt ≤ 271,999Standard synchronous list-price estimate, including historical calls. Promotional through at least 2026-11-21; no confirmed start or expiry is published. The catalog verification date is not a billing boundary.
GPT-5.6 Terra · long contextopenai, openai-codexgpt-5.6-terra$4$0.4$5$18prompt ≥ 272,000
GPT-5.6 Terraopenai, openai-codexgpt-5.6-terra$2$0.2$2.5$12prompt ≤ 271,999
GPT-5.6 Luna · long contextopenai, openai-codexgpt-5.6-luna$0.4$0.04$0.5$1.8prompt ≥ 272,000
GPT-5.6 Lunaopenai, openai-codexgpt-5.6-luna$0.2$0.02$0.25$1.2prompt ≤ 271,999
GPT-5.5 · long contextopenai, openai-codexgpt-5.5$10$1$10$45prompt ≥ 272,000
GPT-5.5openai, openai-codexgpt-5.5$5$0.5$5$30prompt ≤ 271,999
GPT-5.5 Pro · long contextopenai, openai-codexgpt-5.5-pro$60$60$60$270prompt ≥ 272,000No discounted cached-input SKU is published; cached buckets use input price.
GPT-5.5 Proopenai, openai-codexgpt-5.5-pro$30$30$30$180prompt ≤ 271,999No discounted cached-input SKU is published; cached buckets use input price.
GPT-5.4 · long contextopenai, openai-codexgpt-5.4$5$0.5$5$22.5prompt ≥ 272,000
GPT-5.4openai, openai-codexgpt-5.4$2.5$0.25$2.5$15prompt ≤ 271,999
GPT-5.4 Pro · long contextopenai, openai-codexgpt-5.4-pro$60$60$60$270prompt ≥ 272,000No discounted cached-input SKU is published; cached buckets use input price.
GPT-5.4 Proopenai, openai-codexgpt-5.4-pro$30$30$30$180prompt ≤ 271,999No discounted cached-input SKU is published; cached buckets use input price.
GPT-5.4 Miniopenai, openai-codexgpt-5.4-mini$0.75$0.075$0.75$4.5standard
GPT-5.4 Nanoopenai, openai-codexgpt-5.4-nano$0.2$0.02$0.2$1.25standard
GPT-4.1openai, openai-codexgpt-4.1, gpt-4.1-20*$2$0.5$2$8standard
GPT-4.1 Miniopenai, openai-codexgpt-4.1-mini, gpt-4.1-mini-20*$0.4$0.1$0.4$1.6standard
GPT-4.1 Nanoopenai, openai-codexgpt-4.1-nano, gpt-4.1-nano-20*$0.1$0.025$0.1$0.4standard
GPT-5.3 Codexopenai, openai-codexgpt-5.3-codex$1.75$0.175$1.75$14standardStandard Codex tier.
GPT-5openai, openai-codexgpt-5$1.25$0.125$1.25$10standardPrior-generation exact ID retained for existing sessions.
GPT-5 Miniopenai, openai-codexgpt-5-mini$0.25$0.025$0.25$2standardPrior-generation exact ID retained for existing sessions.
GPT-5 Nanoopenai, openai-codexgpt-5-nano$0.05$0.005$0.05$0.4standardPrior-generation exact ID retained for existing sessions.
GPT-4oopenai, openai-codexgpt-4o, gpt-4o-20*$2.5$1.25$2.5$10standard
GPT-4o Miniopenai, openai-codexgpt-4o-mini, gpt-4o-mini-20*$0.15$0.075$0.15$0.6standard
Claude Fable 5anthropicclaude-fable-5*$10$1$12.5$50standardCache-write estimate uses the standard 5-minute cache rate.
Claude Mythos 5anthropicclaude-mythos-5*$10$1$12.5$50standardLimited availability; cache-write estimate uses the 5-minute rate.
Claude Opus 5anthropicclaude-opus-5*$5$0.5$6.25$25standardGlobal standard inference; cache-write estimate uses the 5-minute rate.
Claude Opus 4.8anthropicclaude-opus-4-8*$5$0.5$6.25$25standardGlobal standard inference; cache-write estimate uses the 5-minute rate.
Claude Opus 4.7anthropicclaude-opus-4-7*$5$0.5$6.25$25standardGlobal standard inference; cache-write estimate uses the 5-minute rate.
Claude Opus 4.6anthropicclaude-opus-4-6*$5$0.5$6.25$25standardGlobal standard inference; cache-write estimate uses the 5-minute rate.
Claude Opus 4.5anthropicclaude-opus-4-5*$5$0.5$6.25$25standardGlobal standard inference; cache-write estimate uses the 5-minute rate.
Claude Opus 4.1anthropicclaude-opus-4-1*$15$1.5$18.75$75standardRetired on the first-party API; retained for historical logs.
Claude Opus 4anthropicclaude-opus-4-2025*$15$1.5$18.75$75standardRetired on the first-party API; retained for historical logs.
Claude Sonnet 5anthropicclaude-sonnet-5*$2$0.2$2.5$10standardCache-write estimate uses the 5-minute rate.
Claude Sonnet 4.6anthropicclaude-sonnet-4-6*$3$0.3$3.75$15standardCache-write estimate uses the 5-minute rate.
Claude Sonnet 4.5anthropicclaude-sonnet-4-5*$3$0.3$3.75$15standardCache-write estimate uses the 5-minute rate.
Claude Sonnet 4anthropicclaude-sonnet-4-2025*$3$0.3$3.75$15standardCache-write estimate uses the 5-minute rate.
Claude Haiku 4.5anthropicclaude-haiku-4-5*$1$0.1$1.25$5standardCache-write estimate uses the 5-minute rate.
Claude Haiku 3.5anthropicclaude-3-5-haiku*, claude-haiku-3-5*$0.8$0.08$1$4standardRetired on the first-party API; retained for historical logs.
Gemini 3.7 Flashgoogle, gemini, google-aigemini-3.7-flash$0.75$0.075$0.75$3.75from 2026-08-26T00:00:00.000Z; before 2027-01-01T00:00:00.000ZPromotional through 2026-12-31; cache storage token-hours are excluded.
Gemini 3.7 Flash · 2027 rategoogle, gemini, google-aigemini-3.7-flash$1.5$0.15$1.5$7.5from 2027-01-01T00:00:00.000ZOfficial rate beginning 2027-01-01; cache storage token-hours are excluded.
Gemini 3.6 Flashgoogle, gemini, google-aigemini-3.6-flash$0.75$0.075$0.75$3.75from 2026-08-26T00:00:00.000Z; before 2027-01-01T00:00:00.000ZPromotional through 2026-12-31; cache storage token-hours are excluded.
Gemini 3.6 Flash · 2027 rategoogle, gemini, google-aigemini-3.6-flash$1.5$0.15$1.5$7.5from 2027-01-01T00:00:00.000ZOfficial rate beginning 2027-01-01; cache storage token-hours are excluded.
Gemini 3.5 Flashgoogle, gemini, google-aigemini-3.5-flash$1.5$0.15$1.5$9standardCache storage token-hours are excluded.
Gemini 3.5 Flash-Litegoogle, gemini, google-aigemini-3.5-flash-lite$0.3$0.03$0.3$2.5standardCache storage token-hours are excluded.
Gemini 3.1 Pro Preview · long contextgoogle, gemini, google-aigemini-3.1-pro-preview, gemini-3.1-pro-preview-customtools$4$0.4$4$18prompt ≥ 200,001Cache storage token-hours are excluded.
Gemini 3.1 Pro Previewgoogle, gemini, google-aigemini-3.1-pro-preview, gemini-3.1-pro-preview-customtools$2$0.2$2$12prompt ≤ 200,000Cache storage token-hours are excluded.
Gemini 3.1 Flash-Litegoogle, gemini, google-aigemini-3.1-flash-lite$0.25$0.025$0.25$1.5standardText/image/video rate; audio and cache storage are excluded.
Gemini 2.5 Pro · long contextgoogle, gemini, google-aigemini-2.5-pro$2.5$0.25$2.5$15prompt ≥ 200,001Cache storage token-hours are excluded.
Gemini 2.5 Progoogle, gemini, google-aigemini-2.5-pro$1.25$0.125$1.25$10prompt ≤ 200,000Cache storage token-hours are excluded.
Gemini 2.5 Flashgoogle, gemini, google-aigemini-2.5-flash$0.3$0.03$0.3$2.5standardText/image/video rate; audio and cache storage are excluded.
Gemini 2.5 Flash-Litegoogle, gemini, google-aigemini-2.5-flash-lite$0.1$0.01$0.1$0.4standardText/image/video rate; audio and cache storage are excluded.
DeepSeek V4 Flash · peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-v4-flash$0.44$0.014$0.44$1.32inside listed UTC windows; before 2026-09-10T04:00:00.000Z
DeepSeek V4 Flash · off-peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-v4-flash$0.22$0.007$0.22$0.66outside listed UTC windows; before 2026-09-10T04:00:00.000Z
DeepSeek V4.1 Flash · peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-flash, deepseek-v4-flash, deepseek-v4-flash-vision-exp$0.3$0.006$0.3$1.2inside listed UTC windows; from 2026-09-10T04:00:00.000ZNew Flash rates and legacy Flash aliases from 2026-09-10 04:00 UTC: https://api-docs.deepseek.com/news/news260910/. Current pricing page retains V4 Pro at its own unchanged rates.
DeepSeek V4.1 Flash · off-peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-flash, deepseek-v4-flash, deepseek-v4-flash-vision-exp$0.15$0.003$0.15$0.6outside listed UTC windows; from 2026-09-10T04:00:00.000ZNew Flash rates and legacy Flash aliases from 2026-09-10 04:00 UTC: https://api-docs.deepseek.com/news/news260910/. Current pricing page retains V4 Pro at its own unchanged rates.
DeepSeek V4 Pro · peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-v4-pro$1.32$0.044$1.32$3.96inside listed UTC windows
DeepSeek V4 Pro · off-peakdeepseek, deepseek-api, deepseek-official, eliza/deepseekdeepseek-v4-pro$0.66$0.022$0.66$1.98outside listed UTC windows
DeepSeek Chat legacydeepseekdeepseek-chat$0.28$0.028$0.28$0.42before 2026-07-24T16:00:00.000ZRetired after 2026-07-24 15:59 UTC; retained for historical logs.
DeepSeek Reasoner legacydeepseekdeepseek-reasoner$0.55$0.14$0.55$2.19before 2026-07-24T16:00:00.000ZRetired after 2026-07-24 15:59 UTC; retained for historical logs.
GLM-5.3zai, z-ai, zhipu, bigmodelglm-5.3$1.4$0.26$1.4$4.4standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-5.2zai, z-ai, zhipu, bigmodelglm-5.2$1.4$0.26$1.4$4.4standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-5.1zai, z-ai, zhipu, bigmodelglm-5.1$1.4$0.26$1.4$4.4standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-5zai, z-ai, zhipu, bigmodelglm-5$1$0.2$1$3.2standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-5 Turbozai, z-ai, zhipu, bigmodelglm-5-turbo$1.2$0.24$1.2$4standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.7zai, z-ai, zhipu, bigmodelglm-4.7$0.6$0.11$0.6$2.2standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.7 FlashXzai, z-ai, zhipu, bigmodelglm-4.7-flashx$0.07$0.01$0.07$0.4standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.6zai, z-ai, zhipu, bigmodelglm-4.6$0.6$0.11$0.6$2.2standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.5zai, z-ai, zhipu, bigmodelglm-4.5$0.6$0.11$0.6$2.2standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.5 Xzai, z-ai, zhipu, bigmodelglm-4.5-x$2.2$0.45$2.2$8.9standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.5 Airzai, z-ai, zhipu, bigmodelglm-4.5-air$0.2$0.03$0.2$1.1standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.5 AirXzai, z-ai, zhipu, bigmodelglm-4.5-airx$1.1$0.22$1.1$4.5standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4 32Bzai, z-ai, zhipu, bigmodelglm-4-32b-0414-128k$0.1$0.1$0.1$0.1standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.7 Flashzai, z-ai, zhipu, bigmodelglm-4.7-flash$0$0$0$0standardGlobal Z.AI endpoint; cached-input storage is currently free.
GLM-4.5 Flashzai, z-ai, zhipu, bigmodelglm-4.5-flash$0$0$0$0standardGlobal Z.AI endpoint; cached-input storage is currently free.
Kimi K3kimi, moonshotkimi-k3$3$0.3$3$15standardStandard realtime tier.
Kimi K2.7 Codekimi, moonshotkimi-k2.7-code$0.95$0.19$0.95$4standardStandard realtime tier.
Kimi K2.7 Code Highspeedkimi, moonshotkimi-k2.7-code-highspeed$1.9$0.38$1.9$8standardHigh-speed serving tier.
Kimi K2.6kimi, moonshotkimi-k2.6$0.95$0.16$0.95$4standardStandard realtime tier.
Kimi K2.5kimi, moonshotkimi-k2.5$0.6$0.1$0.6$3before 2026-09-01T00:00:00.000ZScheduled for retirement on 2026-08-31; retained for historical logs.
Moonshot V1 8Kkimi, moonshotmoonshot-v1-8k$0.2$0.2$0.2$2before 2026-09-01T00:00:00.000ZScheduled for retirement on 2026-08-31; no cache discount is published.
Moonshot V1 32Kkimi, moonshotmoonshot-v1-32k$1$1$1$3before 2026-09-01T00:00:00.000ZScheduled for retirement on 2026-08-31; no cache discount is published.
Moonshot V1 128Kkimi, moonshotmoonshot-v1-128k$2$2$2$5before 2026-09-01T00:00:00.000ZScheduled for retirement on 2026-08-31; no cache discount is published.
Grok 4.6 · long contextxaigrok-4.6$4$1$4$12prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.6xaigrok-4.6$2$0.5$2$6prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok Build 0.1 · long contextxaigrok-build-0.1$2$0.4$2$4prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok Build 0.1xaigrok-build-0.1$1$0.2$1$2prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok 4.5 · long contextxaigrok-4.5$4$0.6$4$12prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.5xaigrok-4.5$2$0.3$2$6prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok 4.3 · long contextxaigrok-4.3$2.5$0.4$2.5$5prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.3xaigrok-4.3$1.25$0.2$1.25$2.5prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok 4.20 Reasoning · long contextxaigrok-4.20-0309-reasoning$2.5$0.4$2.5$5prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.20 Reasoningxaigrok-4.20-0309-reasoning$1.25$0.2$1.25$2.5prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok 4.20 Non-Reasoning · long contextxaigrok-4.20-0309-non-reasoning$2.5$0.4$2.5$5prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.20 Non-Reasoningxaigrok-4.20-0309-non-reasoning$1.25$0.2$1.25$2.5prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Grok 4.20 Multi-Agent · long contextxaigrok-4.20-multi-agent-0309$2.5$0.4$2.5$5prompt ≥ 200,000Standard text API; tool-call fees are excluded.
Grok 4.20 Multi-Agentxaigrok-4.20-multi-agent-0309$1.25$0.2$1.25$2.5prompt ≤ 199,999Standard text API; tool-call fees are excluded.
Mistral Medium 3.5mistralmistral-medium-3-5$1.5$1.5$1.5$7.5standardExact cache discount is not published per model; cached buckets conservatively use input price.
Mistral Large 3mistralmistral-large-2512$0.5$0.5$0.5$1.5standardExact cache discount is not published per model; cached buckets conservatively use input price.
Mistral Small 4mistralmistral-small-2603$0.15$0.15$0.15$0.6standardExact cache discount is not published per model; cached buckets conservatively use input price.
Codestralmistralcodestral-2508$0.3$0.3$0.3$0.9standardExact cache discount is not published per model; cached buckets conservatively use input price.
Command Acoherecommand-a-03-2025$2.5$2.5$2.5$10standardCurrent paid production model. No cache discount is published.
Command R7Bcoherecommand-r7b-12-2024$0.0375$0.0375$0.0375$0.15standardPinned paid model. No cache discount is published.
Command Rcoherecommand-r-08-2024$0.15$0.15$0.15$0.6standardPinned paid model. No cache discount is published.
Command R+coherecommand-r-plus-08-2024$2.5$2.5$2.5$10standardPinned paid model. No cache discount is published.
Command legacycoherecommand$1$1$1$2standardDeprecated; retained for historical logs. No cache discount is published.
Command Light legacycoherecommand-light$0.3$0.3$0.3$0.6standardDeprecated; retained for historical logs. No cache discount is published.
Command R legacycoherecommand-r-03-2024$0.5$0.5$0.5$1.5standardDeprecated; retained for historical logs. No cache discount is published.
Command R+ legacycoherecommand-r-plus-04-2024$3$3$3$15standardDeprecated; retained for historical logs. No cache discount is published.
Qwen 3.7 Maxdashscope, alibaba, qwenqwen3.7-max-2026-06-08$2.5$0.5$3.125$7.5prompt ≤ 1,000,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3.7 Plus · long contextdashscope, alibaba, qwenqwen3.7-plus-2026-05-26$1.2$0.24$1.5$4.8prompt ≥ 256,001Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3.7 Plusdashscope, alibaba, qwenqwen3.7-plus-2026-05-26$0.4$0.08$0.5$1.6prompt ≤ 256,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Max · ≤32Kdashscope, alibaba, qwenqwen3-max-2026-01-23$1.2$0.24$1.5$6prompt ≥ 0; prompt ≤ 32,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Max · 32K–128Kdashscope, alibaba, qwenqwen3-max-2026-01-23$2.4$0.48$3$12prompt ≥ 32,001; prompt ≤ 128,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Max · 128K–256Kdashscope, alibaba, qwenqwen3-max-2026-01-23$3$0.6$3.75$15prompt ≥ 128,001; prompt ≤ 256,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Plus · ≤32Kdashscope, alibaba, qwenqwen3-coder-plus-2025-09-23$1$0.2$1.25$5prompt ≥ 0; prompt ≤ 32,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Plus · 32K–128Kdashscope, alibaba, qwenqwen3-coder-plus-2025-09-23$1.8$0.36$2.25$9prompt ≥ 32,001; prompt ≤ 128,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Plus · 128K–256Kdashscope, alibaba, qwenqwen3-coder-plus-2025-09-23$3$0.6$3.75$15prompt ≥ 128,001; prompt ≤ 256,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Plus · 256K–1Mdashscope, alibaba, qwenqwen3-coder-plus-2025-09-23$6$1.2$7.5$60prompt ≥ 256,001; prompt ≤ 1,000,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Flash · ≤32Kdashscope, alibaba, qwenqwen3-coder-flash-2025-07-28$0.3$0.06$0.375$1.5prompt ≥ 0; prompt ≤ 32,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Flash · 32K–128Kdashscope, alibaba, qwenqwen3-coder-flash-2025-07-28$0.5$0.1$0.625$2.5prompt ≥ 32,001; prompt ≤ 128,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Flash · 128K–256Kdashscope, alibaba, qwenqwen3-coder-flash-2025-07-28$0.8$0.16$1$4prompt ≥ 128,001; prompt ≤ 256,000Singapore international list price; cache read uses the implicit-cache rate.
Qwen 3 Coder Flash · 256K–1Mdashscope, alibaba, qwenqwen3-coder-flash-2025-07-28$1.6$0.32$2$9.6prompt ≥ 256,001; prompt ≤ 1,000,000Singapore international list price; cache read uses the implicit-cache rate.
MiniMax M3 · long contextminimaxMiniMax-M3$0.6$0.12$0.6$2.4prompt ≥ 512,001Standard tier; no separate cache-write price is published.
MiniMax M3minimaxMiniMax-M3$0.3$0.06$0.3$1.2prompt ≤ 512,000Standard tier; no separate cache-write price is published.
MiniMax M2.7minimaxMiniMax-M2.7$0.3$0.06$0.375$1.2standard
MiniMax M2.7 HighspeedminimaxMiniMax-M2.7-highspeed$0.6$0.06$0.375$2.4standard
MiniMax M2.5minimaxMiniMax-M2.5$0.3$0.03$0.375$1.2standard
MiniMax M2.5 HighspeedminimaxMiniMax-M2.5-highspeed$0.6$0.03$0.375$2.4standard

Official sources

Accuracy boundaries

  • Route matching is case-insensitive. Only explicit model aliases and narrowly scoped version-suffix globs are included.
  • Prompt length is the sum of uncached input, cache-read, and cache-write buckets reported for a call.
  • DeepSeek V4 peak windows are evaluated from each model-call start timestamp in UTC.
  • Known promotions and retirements use inclusive validFrom / exclusive validTo instants; calls outside them remain unpriced unless a successor rule is published.
  • Anthropic cache writes use the 5-minute rate because Harness usage records do not expose cache TTL.
  • Qwen cache reads use the implicit-cache rate; explicit cache hits can be cheaper.
  • Known provider routes take precedence; custom providers with recognized model names use first-party list-price estimates, not negotiated corporate rates.
  • Unknown model IDs remain unpriced rather than inheriting a broad family wildcard.
  • A custom pricing array replaces the built-in catalog for that plugin instance.