Providers

September 8, 2026 · View on GitHub

providers/get-provider-usage dispatches every selected provider independently and returns one normalized JSON object per provider. Most providers also expose a normalized providers/get-<id>-usage entrypoint, and every provider is covered by providers/get-provider-health. Claude's specialized helper emits KEY=VALUE data for the dispatcher to normalize; pi uses an inline dispatcher envelope plus providers/get-pi-analytics, and hermes computes its dispatcher envelope straight from ~/.hermes/state.db (indexed, milliseconds) with providers/get-hermes-analytics feeding the expanded telemetry card.

This document is the authoritative reference for adapter authors. It records, for every provider:

  • the coverage level the plugin can truthfully report,
  • the credential / local source it reads,
  • the read-only key-validation endpoint (zero token consumption),
  • the quota / balance API if one exists,
  • the subscription plans and billing model,
  • the dashboard URL where usage lives,
  • the current flagship model(s) and recent changelog highlights,
  • and the official documentation source used.

Reviewed 2026-08-26. Provider APIs change; re-check the linked sources before changing an adapter.

Coverage levels

Every provider maps to exactly one coverage level. The level dictates what the widget can truthfully render.

LevelMeaningExample providers
QuotaReal usedPercent + reset window from a protocol/API.codex, copilot, antigravity, openrouter, zai, glm, fireworks (with account ID), commandcode, opencode (Zen mode)
BalanceRemaining prepaid balance / credits in real currency.kimi, deepseek
AnalyticsConsumption counters (requests/tokens/neurons/cost) with no remaining-quota value.cloudflare (GraphQL), 9router, claude (local), pi (local), hermes (local), opencode (local, default), codex (local, alongside its quota)
Auth / configuredValidates credentials with a read-only endpoint when possible; otherwise reports only that a credential is configured and states the limitation. No usage numbers.gemini, mistral, nvidia, qwen, byteplus, groq, cohere, replicate, together, minimax, xai, kilo, ai21
Local runtimeLocal process / installed models.ollama, vertexai (gcloud)
InformationalNo public read-only API at all; the card just links to the dashboard.perplexity, cursor, cline, kiro, warp, amp

Provider kinds

Since 1.11.0 every provider also carries one or more kinds (independent from the coverage level above). The kind drives the card's badge, icon, and accent color:

KindProvidersMeaning
providereverything else (default)Standard API/account provider.
agentpiLocal coding-agent harness; telemetry only, no account quota.
agent,providerhermesDual nature: agent harness and provider front; the card renders as "Agent · Provider".
gateway9routerRouter that aggregates other providers behind one endpoint.
localollamaLocal runtime, no cloud account.

Coverage level (what data can be truthfully reported) and kind (what the thing is) are orthogonal: for example hermes is Analytics coverage with kinds agent,provider.

Coverage matrix

The matrix below summarises the authentication/billing surface for every supported provider. Legend: ✅ yes · ❌ no public API · ⚠️ partial / dashboard-only.

Provider Coverage Read-only key check Quota / balance API Sub. plan PAYG Env var Dashboard Docs source
codex Quota codex app-server ✅ duration-labelled session / weekly windows, including weekly-only responses ✅ Plus / Pro / Team codex login developers.openai.com Codex app-server
claude Quota + local analytics local ~/.claude (or $CLAUDE_CONFIG_DIR) ✅ 5h / 7d / weekly per-model (limits[]) ✅ Pro / Max claude OAuth, optional CLAUDE_CONFIG_DIR ~/.claude local analytics
copilot Quota snapshot copilot_internal/user ✅ premium (AI credits since 2026-06) + overage; chat/completions shown only when not unlimited ✅ Free / Pro / Pro+ / Education / Business / Enterprise COPILOT_GITHUB_TOKEN, GH_TOKEN, or GITHUB_TOKEN (the gh CLI token is tried first) github.com/settings/copilot GitHub Copilot
antigravity Quota (Cloud Code Assist) ✅ local Antigravity OAuth (keyring / IDE session, auto-refreshed) ✅ Gemini / Claude & OpenAI / unknown family quota + reset; optional per-model detail; per-account failures ✅ Antigravity plan ~/.config/Antigravity IDE Antigravity IDE loadCodeAssist + v1internal:fetchAvailableModels on cloudcode-pa.googleapis.com (multi-account)
gemini Auth GET /v1beta/models ❌ dashboard-only ⚠️ tiered PAYG (no flat sub) ✅ prepaid credits → tiers GEMINI_API_KEY aistudio.google.com ai.google.dev
cloudflare Analytics GET /user/tokens/verify ⚠️ GraphQL analytics (not bill) ✅ Workers Paid \$5/mo ✅ \$0.011 / 1k neurons CLOUDFLARE_API_TOKEN dash.cloudflare.com developers.cloudflare.com
mistral Auth GET /v1/models ❌ dashboard-only ✅ Vibe Pro \$14.99 / Team \$24.99 ✅ per token MISTRAL_API_KEY console.mistral.ai docs.mistral.ai
glm / zai Quota GET /api/monitor/usage/quota/limit ✅ per-window % + reset timestamp ✅ GLM Coding Plan \$18–\$160/mo ✅ per token ZAI_API_KEY / GLM_API_KEY (ZHIPU_API_KEY also accepted; precedence depends on the adapter — see GLM / Z.ai) z.ai/manage-apikey docs.z.ai / open.bigmodel.cn
nvidia Auth ⚠️ GET /v1/models is public and cannot validate a key ❌ dashboard-only ⚠️ rate-limited developer trial ✅ NVIDIA Cloud Credits NVIDIA_API_KEY build.nvidia.com/credits docs.api.nvidia.com
minimax Auth GET /v1/models ❌ dashboard-only ✅ Token Plan \$20/\$50/\$120/mo ✅ per token (M3 50% off) MINIMAX_API_KEY platform.minimax.io platform.minimax.io/docs
commandcode Quota GET /provider/v1/models /alpha/billing/credits (5h + weekly) ✅ Go / GOAT / Pro / Max / Team Pro / Provider API ✅ USD credit balance + per-window USD usage COMMAND_CODE_API_KEY or CLI ~/.commandcode/auth.json commandcode.ai/billing Provider API docs
opencode Quota GET /zen/go/v1/models /zen/go/v1/usage (5h + weekly + monthly) ✅ OpenCode Go \$10/mo OPENCODE_API_KEY or CLI ~/.local/share/opencode/auth.json opencode.ai/zen opencode.ai/docs/go
kimi Balance / Quota GET /v1/models /v1/users/me/balance (funds) · /coding/v1/usages (Kimi Code weekly + 5h) ✅ Kimi Code (tempo tiers) ✅ per token (USD/CNY) MOONSHOT_API_KEY, KIMI_API_KEY, or KIMI_CODING_API_KEY platform.kimi.ai platform.kimi.ai/docs
qwen Auth ⚠️ GET /compatible-mode/v1/models ❌ Alibaba billing console ✅ Coding Plan Pro \$50/mo ✅ per token DASHSCOPE_API_KEY modelstudio.console.alibabacloud.com alibabacloud.com/help
xai Auth GET /v1/api-key ❌ dashboard-only (usage.cost_in_usd_ticks per request) ⚠️ prepaid credits only ✅ per token XAI_API_KEY console.x.ai/billing docs.x.ai
kilo Auth ⚠️ GET /api/gateway/models (no-auth) ❌ dashboard-only (402 signal) ✅ Kilo Pass \$19/\$49/\$199/mo ✅ at-provider-cost credits KILO_API_KEY app.kilo.ai kilo.ai/docs
kiro Informational ❌ no public API ❌ no public API ✅ Free / Pro / Pro+ / Pro Max / Power ⚠️ overage \$0.04/credit app.kiro.dev kiro.dev/docs
9router Local analytics local SQLite/JSON ✅ local requests/tokens/cost ~/.9router local store
pi Local analytics local JSONL session logs ✅ local cost/tokens — no quota API (pi has no rate limits) local store
hermes Local analytics local SQLite state (~/.hermes/state.db) ✅ local sessions/tokens/cost per model and project — the agent side; provider billing (Nous Portal / routed providers) stays on the portal NousResearch/hermes-agent
openrouter Quota/balance GET /api/v1/key ✅ limit/daily/monthly/balance ✅ per token OPENROUTER_API_KEY openrouter.ai/credits openrouter.ai/docs
deepseek Balance balance endpoint ✅ account + granted balance ✅ per token DEEPSEEK_API_KEY platform.deepseek.com api-docs.deepseek.com
ollama Local runtime GET /api/tags ✅ installed + running models OLLAMA_HOST localhost:11434 docs.ollama.com
vertexai Local runtime gcloud auth ⚠️ GetUsage exists with separate IAM Signature V4 credentials gcloud console.cloud.google.com cloud.google.com/vertex-ai
byteplus Auth GET /api/v3/models ❌ dashboard-only BYTEPLUS_API_KEY console.byteplus.com docs.byteplus.com
together Auth GET /v1/models ❌ no documented read-only credits endpoint TOGETHER_API_KEY api.together.ai docs.together.ai
groq Auth GET /v1/models ❌ dashboard-only GROQ_API_KEY console.groq.com console.groq.com/docs
cohere Auth ✅ models API ❌ dashboard-only COHERE_API_KEY dashboard.cohere.com docs.cohere.com
replicate Auth ✅ account endpoint ❌ dashboard-only ✅ per second REPLICATE_API_TOKEN replicate.com/account replicate.com/docs
fireworks Quota (with account ID) ✅ inference models API GET /v1/accounts/{account_id}/quotas FIREWORKS_API_KEY; optional FIREWORKS_ACCOUNT_ID app.fireworks.ai docs.fireworks.ai
ai21 Auth GET /studio/v1/models ❌ dashboard-only AI21_API_KEY studio.ai21.com docs.ai21.com
perplexity Informational ✅ Pro perplexity.ai/settings docs.perplexity.ai
cursor Informational ✅ Hobby / Pro / Business cursor.com/settings cursor.com
cline Informational app.cline.bot cline.bot
warp Informational ✅ Warp Pro app.warp.dev warp.dev
amp Informational ✅ Amp Pro ampcode.com ampcode.com

Local fallbacks and aliases

OpenRouter → 9Router local fallback. When OPENROUTER_API_KEY is unset, the dispatcher serves the card from the local 9Router store instead (fetch_9router_native openrouter openrouter soft); the account label reads "openrouter via 9router (local)" so locally routed data is never mistaken for upstream OpenRouter API quota. With the variable set, the real /api/v1/key snapshot is used.

Aliases. The dispatcher (get-provider-usage) and health checker (get-provider-health) accept these case-insensitive aliases:

CanonicalAliases
commandcodecmd, cmdcode
antigravityagy
kimimoonshot
glmzhipu
zaiz.ai
nvidianim
vertexaivertex
byteplusark, modelark
qwendashscope, alibaba
xaigrok

Provider reference

Detailed adapter notes for the focus providers (Gemini, Cloudflare, Mistral, GLM/Z.ai, NVIDIA, MiniMax, Kimi, Qwen, xAI, Kilo, Kiro). All HTTP probes below are read-only and consume no tokens.

Gemini (Google)

API basehttps://generativelanguage.googleapis.com/v1beta (OpenAI-compat: /v1beta/openai/). No China mirror; enterprise path is Vertex AI.
Env varGEMINI_API_KEY (fallbacks GOOGLE_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY); also detected via local Gemini CLI (~/.gemini/oauth_creds.json).
AuthHeader x-goog-api-key: <key> (preferred) or query ?key=<key>.
Key checkGET /v1beta/models200 lists models; 403 missing key; 400 malformed. Zero tokens.
Quota / balance❌ None. Usage, spend, and rate limits are dashboard-only.
PlansFree + tiered PAYG (Tier 1 $250 cap → Tier 3 $20k+ cap, auto-upgrade by spend/age). No flat monthly API subscription. Enterprise = Gemini Enterprise Agent Platform.
BillingPrepaid credits → pay-as-you-go per token. Batch API 50% off; cached/prompt discounts.
Flagship pricing (/M tok)gemini-3.5-flash $1.50/$9.00 · gemini-3.1-pro-preview $2.00/$12.00 (≤200k) · gemini-3.1-flash-lite $0.25/$1.50 · gemini-2.5-pro $1.25/$10.00.
Dashboardaistudio.google.com (rate limits), console.cloud.google.com/billing.
ChangelogGemini 3.5 Flash GA; Gemini 3.1 Pro preview (vibe-coding/customtools). Gemini 2.0 Flash/Flash-Lite shut down 2026-06-01. Imagen 4 deprecated (→ Gemini 2.5 Flash Image). Veo 3/2 → Veo 3.1. Gemini Embedding 2 (multimodal).
Adapterfetch_gemini_native — API-key probe; falls back to local Gemini CLI OAuth detection.

Cloudflare (Workers AI)

API basehttps://api.cloudflare.com/client/v4 (account-scoped AI: /accounts/{account_id}/ai/...; OpenAI-compat: /accounts/{account_id}/ai/v1).
Env varCLOUDFLARE_AI_TOKEN (preferred) or CLOUDFLARE_API_TOKEN; CLOUDFLARE_ACCOUNT_ID unlocks GraphQL analytics.
AuthAuthorization: Bearer <token>. Token needs Workers AI - Read/Edit.
Key checkGET /user/tokens/verify200 {"success":true,"result":{"status":"active"}}. Zero neurons.
Quota / balance⚠️ GraphQL analytics only (not the bill). POST /graphql with aiInferenceAdaptiveGroups dataset scoped by accountTag. Returns sum { requests neurons } and dimensions { date modelName taskType }. The dedicated tutorial page was removed in 2025 — verify node/fields via introspection before relying on it. No REST "remaining neurons" endpoint.
PlansWorkers Free: 10,000 neurons/day. Workers Paid ($5/mo): same free allotment then $0.011/1k neurons. Enterprise: custom.
BillingPer neuron ($0.011/1k). Per-model rates also expressed as /M tokens (e.g. `@cf/zai-org/glm-5.2` \1.40/$4.40).
Dashboarddash.cloudflare.com → account → Workers AI.
Changelog2026-06-16 GLM-5.2; 2026-06-12 Kimi K2.7 Code; 2026-04-20 Kimi K2.6 (reasoning field, chat_template_kwargs.thinking); 2026-03-19 prompt caching (x-session-affinity); 2026-02-13 GLM-4.7-Flash; 2025-08-05 gpt-oss-120b/20b; 2024-05-17 OpenAI-compat added; unit-based neuron pricing since 2024-09. Major deprecation wave 2026-05-30 (Llama-2/3/3.1 base+AWQ, Mistral 7B, Gemma 7B/3-12B, phi-2).
Adapterfetch_cloudflare_native — token verify; if CLOUDFLARE_ACCOUNT_ID is set, queries aiInferenceAdaptiveGroups for 7-day totals + latest day (graceful fallback to the token-verified note card).

Mistral AI

API basehttps://api.mistral.ai/v1. No regional mirror.
Env varMISTRAL_API_KEY.
AuthAuthorization: Bearer <key>.
Key checkGET /v1/models200; 401 on bad key. Zero tokens. (/v1/billing and /v1/usage return 404.)
Quota / balance❌ None. Dashboard-only ("Mistral Studio dashboard offers detailed tracking of API usage").
PlansVibe consumer/coding subs: Free / Pro $14.99 / Team $24.99 / Enterprise / Education $5.99. API = PAYG per token. Batch 50% off; cached input 10% of input price.
BillingPer token (per 1M): mistral-medium-3.5 $1.5/$7.5 (coding flagship) · mistral-large-3 $0.5/$1.5 · mistral-small-4 $0.1/$0.3 · devstral-2 $0.4/$2.0 (agentic coding) · codestral $0.3/$0.9 (FIM) · magistral-medium $2.0/$5.0.
Dashboardconsole.mistral.ai (keys, usage, billing). Pricing: mistral.ai/pricing.
Changelog2026-04 Mistral Medium 3.5; 2026-03 Mistral Small 4 (Apache 2.0); product restructure into Vibe / Studio / Admin (La Plateforme rebranded to Studio). New Workflows, Connectors, Observability. Fine-tuning + old Agents endpoints deprecated.
Adapterfetch_mistral_native/v1/models validation only.

GLM / Z.ai (Zhipu AI)

API baseGlobal https://api.z.ai/api/paas/v4; Coding-Plan https://api.z.ai/api/coding/paas/v4; China https://open.bigmodel.cn/api/paas/v4. Fully OpenAI-compatible.
Env varPer adapter: fetch_zai_native (global card) uses ZAI_API_KEYGLM_API_KEYZHIPU_API_KEY; fetch_glm_native (China console card) uses GLM_API_KEYZHIPU_API_KEYZAI_API_KEY. Any one key drives both cards.
AuthAuthorization: Bearer <key>.
Key checkGET /api/monitor/usage/quota/limit200 with success: true and data.limits[]; 401/403 on bad key. Zero tokens. Falls back to GET /paas/v4/models when quota endpoint is unavailable.
Quota / balance/api/monitor/usage/quota/limit returns data.limits[] — each limit has type (TIME_LIMIT or TOKENS_LIMIT), percentage (0–100), nextResetTime (epoch ms), remaining, unit, and number. Live cross-check against the Z.ai Usage page confirms the period unit table used by the adapter: unit=4 → 5-hour session window, unit=6 → weekly quota, unit=5 → monthly web/search/reader quota, unit=3 → total token allotment (no reset window). Also returns data.level (plan tier, e.g. lite).
PlansPAYG per token or GLM Coding Plan (Lite $18, Pro $72, Max $160/mo; quarterly/annual discounts). Subscription usage exposes a 5-hour session quota, weekly token quota, and monthly Web Search / Reader / Zread quota; supported in Claude Code, Cline, OpenCode, Roo, Kilo, Crush, Goose, OpenClaw. GLM-5.2 & GLM-5-Turbo consume 3× quota at peak (14:00–18:00 UTC+8), 2× off-peak (1× off-peak promo through end of September).
BillingPer 1M tokens: glm-5.2 $1.40/$4.40 (cached $0.26) · glm-5/glm-5-turbo $1.0–$1.2/$3.2–$4.0 · glm-4.7/4.6/4.5 $0.60/$2.20 · glm-4.7-flash & glm-4.5-flash Free · glm-5v-turbo (vision) $1.2/$4.0. Web Search $0.01/use.
DashboardAPI keys, subscription (Coding Plan), and billing (finance). China: open.bigmodel.cn.
ChangelogGLM-5.2 live (1M lossless context, 128K max output, MCP, structured output). Lineage: 4.5 → 4.6 → 4.7 → 5 → 5-Turbo → 5.1 → 5.2. New GLM-5V-Turbo (vision coding), GLM-Image, CogVideoX-3, GLM-ASR-2512, GLM-OCR. Coding Plan restructured (legacy plans migrated by 2026-04-30).
Adapterfetch_glm_native (China console, open.bigmodel.cn) and fetch_zai_native (global, api.z.ai). Both call GET /api/monitor/usage/quota/limit for real quota data; fall back to GET /models auth-only check if the quota endpoint is unavailable. Sorts limits by urgency (highest % first among timed windows), then maps to primary / secondary / tertiary.

NVIDIA (NIM / build.nvidia.com)

API basehttps://integrate.api.nvidia.com/v1 (OpenAI-compatible). No regional mirror; backed by DGX Cloud.
Env varNVIDIA_API_KEY (NGC personal key, nvapi- prefix).
AuthAuthorization: Bearer <key>.
Key checkGET /v1/models is a public catalog and cannot validate a key; the card truthfully reports only that a key is set.
Quota / balance❌ None documented. NVCF/Cloud Functions billing doc is auth-gated (401). Balance is dashboard-only. Practical exhaustion signal: inference returns 402/429.
PlansRate-limited Developer trial, NVIDIA Cloud Credits (USD blocks), and NVIDIA AI Enterprise for self-hosted NIM.
BillingCredit-based per request (1 credit ≈ 1 standard request; heavy models cost more). No posted $/M-token table for the hosted catalog.
Dashboardbuild.nvidia.com/credits. NGC keys at org.ngc.nvidia.com.
ChangelogNemotron-3 family (Ultra-550B, Super-120B-A12B, Nano-30B-A3B, Nano-Omni-30B). Heavy expansion into third-party frontier models (DeepSeek V4 Pro/Flash, Mistral Large 3 675B + Medium 3.5 + Small 4, MiniMax M2.7/M3, Stepfun, Z.ai GLM 5.1/4.7, Qwen3-Coder-480B). New async 202+polling pattern for heavy endpoints. Healthcare microservices (AlphaFold2, Boltz2, OpenFold3).
Adapterfetch_nvidia_native — public /v1/models catalogue probe; it does not claim that the key is validated.

MiniMax

API basehttps://api.minimax.io/v1 (OpenAI-compat); Anthropic-compat https://api.minimax.io/anthropic. Legacy TTS host api.minimax.chat. No .cn host.
Env varMINIMAX_API_KEY. Two key types: pay-as-you-go API Key vs Token-Plan Subscription Key (not interchangeable).
AuthAuthorization: Bearer <key>.
Key checkGET /v1/models200 lists MiniMax-M3, MiniMax-M2.7, MiniMax-M2.5… Zero tokens.
Quota / balance❌ None. Dashboard-only. Token-Plan usage shown as a console bar (5-hour rolling + weekly). Errors: 1004 auth failed, 1008 insufficient balance, 1002 rate limit, 1039 token limit exceeded.
PlansToken Plan (replaces old "Coding Plan"): Plus $20 / Max $50 / Ultra $120/mo — full-spectrum multimodal, 5h+weekly windows, no rollover. Credits: $5/$25/$100 packs (1000cr = $1, 365-day). PAYG also available.
BillingPer 1M tokens: MiniMax-M3 (≤512K, 50% off) $0.30/$1.20 (cache $0.06) · >512K $0.60/$2.40 · Priority tier 1.5× · MiniMax-M2.7 $0.30/$1.20 · MiniMax-M2.7-highspeed $0.60/$2.40. Audio speech-2.8-hd $100/M chars; Hailuo video $0.19–$0.56/clip.
Dashboardplatform.minimax.io: keys /user-center/basic-information/interface-key, balance /user-center/payment/balance, Token Plan /user-center/payment/token-plan.
Changelog2026-06-01 MiniMax-M3 (1M ctx, adaptive thinking, coding SOTA). 2026-03-18 M2.7/M2.7-highspeed. 2026-02 M2.5. 2025-12-22 M2.1. 2025-10-27 M2 + Hailuo-2.3. Token Plan replaced Coding Plan (broader coverage, separate Subscription Key).
Adapterfetch_minimax_native/v1/models validation.

Command Code

API basehttps://api.commandcode.ai/provider/v1 (OpenAI/Anthropic-compatible); quota https://api.commandcode.ai/alpha/billing/credits (Bearer, no cookies).
CredentialsCOMMAND_CODE_API_KEY (preferred when available to DMS), or the apiKey in CLI-owned ~/.commandcode/auth.json after cmd login. Both are the same Provider API key used for /provider/v1/models, /alpha/whoami, /alpha/billing/credits.
AuthAuthorization: Bearer ***
Quota / balance/alpha/billing/credits returns windowLimits.fiveHour (5h USD used/cap, Unix-ms resetAt), windowLimits.weekly (7-day USD used/cap, Unix-ms resetAt), and credits.monthlyCredits (remaining USD this billing cycle). /alpha/billing/subscriptions exposes planId (e.g. individual-goat) and currentPeriodEnd.
PlansGo $1 / $10 credit · GOAT $10 / $70 credit, 5h $14 / weekly $35 · Pro $20 / $80 credit, 5h $16 / weekly $40 · Max 10× $100 / $150 credit, 5h $45 / weekly $90 · Max 20× $200 / $300 credit, 5h $90 / weekly $180 · Team Pro $40 / $40 credit, 5h $12 / weekly $24 · Provider API $15 + pay-as-you-go.
BillingPer-token, no markup on the Provider plan. Credits roll over and never expire. Window caps are USD ceilings — model deals (e.g. minimax-m3 2×, google/gemini-3.7-flash 50% off, mimo-v2.5-pro 99% off) apply automatically.
StabilityThe /alpha/ namespace is experimental — no documented versioned contract, can change without notice. Adapter degrades to the documented /provider/v1/models endpoint on alpha failure and emits a clearly-labeled "quota endpoint unavailable" note; it never fabricates a percentage.
Dashboardcommandcode.ai/billing — billing, plan, and credit top-ups.
Adapterfetch_commandcode_native/alpha/billing/credits + /alpha/whoami + /alpha/billing/subscriptions, with documented-endpoint fallback.

OpenCode

Two independent modes. The local mode is the default whenever an OpenCode database exists, because OpenCode is commonly driven against third-party and free-tier providers with no Zen subscription — the Zen quota endpoint would only report an entitlement error there. Set OPENCODE_USAGE_SOURCE=api to force the Zen path.

Local source${OPENCODE_DATA_DIR:-${XDG_DATA_HOME:-$HOME/.local/share}/opencode}/opencode.db (SQLite, read-only). Assistant messages carry providerID, modelID, tokens and cost; only usage metadata is read, never message content. Costs come from whatever OpenCode recorded — a missing cost makes the window unknown () rather than $0.
Local adapterfetch_opencode_native (local branch) + providers/get-local-analytics opencode, cached 120s.

OpenCode Go (Zen quota mode)

API basehttps://opencode.ai/zen/go/v1 — OpenCode Zen's Go-plan quota surface, alongside the OpenAI-compatible /zen/v1 inference API.
CredentialsOPENCODE_API_KEY (preferred when available to DMS), or the key in CLI-owned ${XDG_DATA_HOME:-$HOME/.local/share}/opencode/auth.json (saved by opencode auth login). The adapter follows this XDG path on Linux, the plugin's supported platform. The same key created in OpenCode Studio/Zen drives both inference and quota introspection.
AuthAuthorization: Bearer ***
Quota / balance/zen/go/v1/usage returns rollingUsage (5h window), weeklyUsage, and monthlyUsage, each {status, resetInSec, usagePercent} — percentages are computed server-side, so the adapter never derives them from raw byte/token counts. The adapter accepts a quota response only when all three windows have valid percentage/reset values; otherwise it falls back to auth-only status. resetInSec is seconds-from-now, converted to an absolute ISO 8601 timestamp. useBalance: true is rendered as balance fallback enabled; the endpoint does not return the balance amount, so none is invented.
PlansOpenCode Go $10/mo — $12 of usage per 5-hour window, $30 weekly, $60 monthly.
BillingFlat monthly subscription; no PAYG surfaced through this endpoint.
StabilityThe /zen/go/ namespace shipped alongside the Go plan (opencode PR #16513) and has no documented versioned contract. Adapter degrades to /zen/go/v1/models (auth-only) on failure and emits a clearly-labeled "quota endpoint unavailable" note; it never fabricates a percentage. A missing/invalid key returns 401 AuthError; a valid key without a Go subscription returns 403 EntitlementError — both surface as hard errors, not soft notes.
Dashboardopencode.ai/zen — Zen usage and Go plan management.
Adapterfetch_opencode_native (API branch) — /zen/go/v1/usage, with /zen/go/v1/models fallback.

Kimi (Moonshot AI)

Two systemsKimi exposes two independent quota surfaces with non-interchangeable keys.Open Platform (prepaid balance): sk-xxx key on api.moonshot.ai/.cn. ② Kimi Code / Coding Plan (subscription quota): sk-kimi-xxx key on api.kimi.com/coding/v1 — the same data the Kimi CLI /usage command reads. The adapter routes by key (see Env var / Adapter).
API baseOpen Platform — Global https://api.moonshot.ai/v1 (USD); China https://api.moonshot.cn/v1 (CNY). Coding Plan — https://api.kimi.com/coding/v1 (override with KIMI_BASE_URL). Platform rebranded: platform.moonshot.aiplatform.kimi.ai (global) / platform.moonshot.cnplatform.kimi.com; API hosts unchanged.
Env varBalance: MOONSHOT_API_KEY (fallback KIMI_API_KEY); optional MOONSHOT_API_BASE overrides host probing. Coding Plan: KIMI_CODING_API_KEY, or a KIMI_API_KEY/MOONSHOT_API_KEY carrying the sk-kimi- prefix; optional KIMI_BASE_URL. An explicit coding key wins over the balance path.
AuthAuthorization: Bearer <key> (both systems). Coding Plan requests also send User-Agent: KimiCLI/1.6.
Key checkBalance: GET /v1/models200; 401 on bad key. Coding Plan: the /usages probe doubles as the key check (401/403 = wrong key type, 404 = wrong base URL).
Quota / balanceBalance: ✅ GET /v1/users/me/balance{data:{available_balance, voucher_balance, cash_balance}} (USD on .ai, CNY on .cn; available_balance = cash + voucher). Coding Plan: ✅ GET /coding/v1/usages (older deployments answer on /usage) → weekly + 5-hour windows with used/limit/remaining and a reset time; mapped weekly → primary, 5h → secondary.
PlansOpen Platform: no subscription — PAYG + prepaid top-up vouchers; rate-limit tiers scale with cumulative recharge ($1 Tier0 → $3,000 Tier5). Kimi Code subscription tiers (named after musical tempos): Adagio $0 (no Kimi Code), Moderato $19, Allegretto $39, Allegro $99, Vivace $199/mo (annual ≈ −20%). Quota refreshes on a 7-day cycle plus a 5-hour burst limiter (≈300–1,200 requests/5h, up to 30 concurrent by tier); entry paid tier ≈ 2,048 Kimi Code requests/week. 403 "reached your usage limit for this billing cycle" = weekly quota exhausted.
BillingPer 1M tokens (256K ctx): kimi-k2.7-code $0.95/$4.00 (cache hit $0.19) · kimi-k2.7-code-highspeed $1.90/$8.00 · kimi-k2.6 $0.95/$4.00 · kimi-k2.5 $0.60/$3.00. Automatic context caching.
DashboardBalance: platform.kimi.ai/console (global) / platform.kimi.com/console (China). Coding Plan: Kimi Code console + membership page; in-CLI /usage.
Changelogkimi-k3 (max quality) and kimi-k2.7-code / -highspeed (routine coding, ~180 tok/s). kimi-k2.6 (multimodal flagship). kimi-k2.5 + moonshot-v1 retire 2026-08-31 → migrate to kimi-k2.7-code or kimi-k3. kimi-k2 series deprecated 2026-05-25; kimi-latest removed 2026-01-28; kimi-thinking-preview removed 2025-11-11. Full OpenAPI spec at platform.kimi.ai/docs/openapi.json.
Adapterfetch_kimi_native routes by key: sk-kimi-/KIMI_CODING_API_KEYfetch_kimi_code_native (Coding Plan GET /usages → weekly primary + 5h secondary, source: kimi-code); otherwise the balance API with global→China host probing (source: kimi-api, primary = available balance, secondary = voucher/cash split).

Qwen / DashScope (Alibaba Model Studio)

API base5 regional modes: International (Singapore) https://dashscope-intl.aliyuncs.com/compatible-mode/v1; US (Virginia) https://dashscope-us.aliyuncs.com/compatible-mode/v1; China (Beijing) https://dashscope.aliyuncs.com/compatible-mode/v1; HK https://cn-hongkong.dashscope.aliyuncs.com/compatible-mode/v1; EU (Frankfurt) https://{workspace}.eu-central-1.maas.aliyuncs.com/compatible-mode/v1. Coding Plan uses https://coding-intl.dashscope.aliyuncs.com/v1 with a dedicated sk-sp-… key.
Env varDASHSCOPE_API_KEY (fallback QWEN_API_KEY); optional DASHSCOPE_WORKSPACE_ID header. Coding Plan uses sk-sp-….
AuthAuthorization: Bearer <key> + optional X-DashScope-WorkSpace.
Key check⚠️ GET /compatible-mode/v1/models is used by the adapter but is not officially documented in the OpenAI-compat surface — treat as best-effort. 401 confirms a bad key.
Quota / balance❌ None. Billing flows through Alibaba Cloud User Center (per-minute inference, per-hour batch; monthly settle).
PlansCoding Plan Pro $50/mo: 6,000 req/5h, 45,000/week, 90,000/month — qwen3.5-plus, kimi-k2.5, glm-5, MiniMax-M2.5, qwen3-coder-next/+. Lite plan closed to new subs since 2026-03-20. Free quota (1M in+out tokens × 90 days) on International region only.
BillingPer 1M tokens (intl, USD): qwen3-max $1.20/$6.00 (≤32K; doubles >128K) · qwen3.5-plus $0.40/$2.40 (1M ctx) · qwen3.5-flash $0.10/$0.40 · qwq-plus $0.80/$2.40. Batch 50% off; Context Cache (input-only). CNY on China-mainland.
Dashboardmodelstudio.console.alibabacloud.com (intl); bailian.console.alibabacloud.com (China). Billing: Alibaba Cloud User Center.
Changelog2026-02 qwen3.5-plus-2026-02-15, qwen3.5-flash-2026-02-23; 2026-01-23 qwen3-max-2026-01-23 (web-search/code-interpreter in thinking mode). New open-source qwen3.5-{397b-a17b,120b-a10b,27b,35b-a3b}. Qwen-Turbo deprecated (→ Qwen-Flash). Coding Plan Lite closed to new subs.
Adapterfetch_qwen_native/compatible-mode/v1/models validation.

xAI (Grok)

API basehttps://api.x.ai/v1 (inference, OpenAI-compat). Key management: https://management-api.x.ai (separate, June 2025).
Env varXAI_API_KEY.
AuthAuthorization: Bearer <key>.
Key checkGET /v1/api-key200 returns key metadata: {redacted_api_key, name, user_id, team_id, api_key_id, api_key_blocked, api_key_disabled, team_blocked, acls, create_time, modify_time}. 401 without key. This is the best validator — it also surfaces blocked/disabled state. (GET /v1/models also works and returns per-token prices.)
Quota / balance❌ No remaining-credits field on /v1/api-key. Account balance is dashboard-only. Per-request cost is in every response's usage.cost_in_usd_ticks (1 USD = 1e10 ticks) and usage.cost_in_nano_usd (Responses API); enable via stream_options.include_usage:true.
PlansAPI = prepaid credits (no named tiers). service_tier:"default" vs "priority" (2× billing, higher priority — June 2026). Consumer Grok subs (Free/SuperGrok/SuperGrok Heavy) are separate from api.x.ai.
BillingPer 1M tokens: grok-4.3 (flagship, 1M ctx) $1.25/$2.50 · grok-4.20-0309-{reasoning,non-reasoning,multi-agent} $1.25/$2.50 · grok-build-0.1 (coding, aliases grok-code-fast-1, grok-code-fast, 256K ctx) $1.00/$2.00 (cache $0.20). Batch 20–50% off. Tools billed per 1k calls (web/x search $5, code exec $5, attachment $10). Image $0.02–$0.05/img; video $0.05–$0.08/sec.
Dashboardconsole.x.ai: keys /team/default/api-keys, billing /billing, models /team/default/models. Status: status.x.ai.
Changelog2026-06 Priority Processing + Files public URLs. 2026-05 grok-build-0.1 coding model + Grok Build CLI + Context Compaction API + WebSocket Responses. 2026-04 cost_in_usd_ticks on every response. 2026-03 Grok 4.20 + Multi-agent. 2025-12 Voice Agent API GA. 2025-11 Grok 4.1 Fast + Files API GA + Remote MCP. 2025-10 agentic server-side tools GA. 2025-09 Responses API GA. 2025-07 Grok 4. Anthropic-compat /v1/messages & /v1/complete deprecated; max_tokensmax_completion_tokens.
Adapterfetch_xai_nativeGET /v1/api-key; surfaces key name + blocked/disabled flags in the note.

Kilo (kilo.ai)

API baseGateway https://api.kilo.ai/api/gateway (OpenAI-compat; model string is provider/model, e.g. anthropic/claude-sonnet-4.5).
Env varKILO_API_KEY (JWT tied to the Kilo account). Optional headers X-KiloCode-OrganizationId, X-KiloCode-TaskId, X-KiloCode-Version.
AuthAuthorization: Bearer <key>. Anonymous access allowed only for :free models (rate-limited per IP).
Key check⚠️ GET /api/gateway/models is documented as no-auth — a 200 is inconclusive; a 401 indicates a malformed key. No dedicated /api-key validate endpoint. The reliable bad-key signal is a 402 on a paid call ({error:{code:402, metadata:{buyCreditsUrl}}}).
Quota / balance❌ None. Per-request usage.cost_microdollars (1 USD = 1e6 microdollars), input_tokens, output_tokens, cache_*_tokens, is_byok. Balance is dashboard-only.
PlansKilo Code agent: Free/OSS, Teams $15/user/mo, Enterprise. Inference: Auto Free / BYOK / local ($0); Kilo Gateway ($0 + PAYG at exact provider rates, zero markup); Kilo Pass — Starter $19 (≈$26.60), Pro $49 (≈$68.60), Expert $199 (≈$278.60) with up to 50% bonus credits. KiloClaw managed OpenClaw $55/mo.
BillingCredit-based, exact upstream provider rates, zero markup (microdollar precision). Free models (:free) cost nothing (200 req/hour/IP → 429). BYOK = $0 on Kilo's side.
Dashboardapp.kilo.ai: /credits, /subscriptions, /claw. Docs: kilo.ai/docs. Status: status.kilo.ai.
ChangelogAgent rebuilt on Kilo CLI. Auto Model router (Frontier/Balanced/Free). Kilo Pass subscription with bonus credits. KiloClaw managed hosting. Kilo Marketplace (spend credits on partner plans; MiniMax first). New BYOK: Inceptron, Kimi Code, Z.ai Coding Plan, Xiaomi Token Plans. Routes 500+ models across 60+ providers (no first-party model).
Adapterfetch_kilo_native — best-effort /api/gateway/models probe; a 401 rejects the key, otherwise falls back to the configured-status note.

Kiro (kiro.dev — AWS)

API baseNone. Kiro is a subscription-only agentic IDE/CLI/Web with no public REST inference API and no API-key model. Login is interactive SSO (GitHub / Google / AWS Builder ID / AWS IAM Identity Center). CLI install: curl -fsSL https://cli.kiro.dev/install | bash.
Env var— (no KIRO_API_KEY).
AuthSSO login (not an API key). Enterprise uses AWS IAM Identity Center / SAML/SCIM.
Key check❌ None.
Quota / balance❌ None. Credits surfaced in-product only (IDE/CLI/Web subscription dashboard), refreshed at least every 5 minutes.
PlansFree $0 (50 cr) · Pro $20 (1,000 cr) · Pro+ $40 (2,000 cr) · Pro Max $100 (5,000 cr, new 2026-06-10) · Power $200 (10,000 cr). $20 sign-up bonus on first upgrade. No daily/weekly rate limits. Credits metered to 0.01, no rollover, processed on the 1st. Overage $0.04/credit (opt-in). GovCloud ~+20%, no Free. HIPAA eligible (IDE+CLI) since 2026-05-26.
BillingMonthly sub + credit multiplier per model: Auto 1×, Claude Sonnet 4.6 1.3×, Claude Opus 4.8 2.2×, Haiku 4.5 0.4×, DeepSeek v3.2 0.25×, MiniMax M2.5 0.25×. Taxes not included.
Dashboardapp.kiro.dev/settings/account. Docs: kiro.dev/docs/billing. Pricing: kiro.dev/pricing.
Changelog2026-06-17 CLI v2.8 (CLI v3 early access). 2026-06-12 CLI v2.7 (/goal loops, queue steering). 2026-06-10 Pro Max tier. 2026-05-29 Claude Opus 4.8 (2.2×, 1M ctx). 2026-05-26 HIPAA eligible. CLI v2.6/v2.7 transcript export, persistent prefs.
AdapterInformational card only (json_note_usage kiro-local). Links to app.kiro.dev. No scriptable surface.

Codex local telemetry

Alongside the app-server quota windows, the Codex card shows a local telemetry panel built from ~/.codex/sessions/**/*.jsonl (and archived_sessions/), honouring CODEX_HOME.

Sourceevent_msg / token_count events. Only usage counters, cwd, and the model name are read — never prompt or response content.
Cumulative counterstotal_token_usage is cumulative for the session and is re-emitted unchanged on turns that add no tokens. The adapter sums positive deltas between consecutive snapshots, so repeated values are not double-counted and a counter reset (a new total lower than the previous one) restarts from zero instead of producing a negative row. last_token_usage is deliberately unused: it is also repeated.
CostCodex sessions record no cost, so every window reports cost as unknown (). Consumption against a Plus/Pro subscription is not a dollar charge, and none is invented.
BucketingRows are bucketed by the event's own local calendar day (via TZ), unlike pi/Hermes which bucket a whole session to its start day.
Adapterproviders/get-local-analytics codex, cached 120s.

Codex backend launch discipline

Every codex app-server startup runs a marketplace refresh round; when an upgrade keeps being interrupted, each round leaves git temp directories under ~/.codex/.tmp/ that Codex itself does not clean up (upstream issues: openai/codex#30620, #36093). A quota poll must therefore launch as few backends as possible:

  • Daemon/proxy mode. When the Codex CLI supports it (standalone-installer installs), get-codex-usage runs codex app-server daemon start (idempotent) and speaks each poll through codex app-server proxy, so steady-state usage never starts a backend at all. npm/brew/distro installs without the daemon subcommand fall back to spawning a backend per refresh.
  • CODEX_APP_SERVER_MODE=spawn forces the spawn path and skips the daemon probe.
  • Single-flight. A flock on ~/.cache/AiOverviewControl/codex-usage.lock serializes launches; an invocation arriving while another refresh runs serves the cached snapshot, or waits up to CODEX_LOCK_WAIT seconds (default 8) before returning a structured error.
  • Freshness gate. A cached snapshot younger than CODEX_FRESH_TTL seconds (default 60) is answered directly — widget reload bursts no longer spawn anything.
  • Graceful teardown. The backend exits on its own when stdin closes; the adapter waits briefly for that self-exit before falling back to SIGTERM/SIGKILL, so startup work in flight is not orphaned mid-clone.

To reclaim disk from already-leaked staging clones (safe when no Codex session is running):

find ~/.codex/.tmp/marketplaces/.staging -mindepth 1 -maxdepth 1 \
  -type d -name 'marketplace-upgrade-*' -exec rm -rf {} +
find ~/.codex/.tmp -mindepth 1 -maxdepth 1 -type d -name 'git-*' -exec rm -rf {} +

Utility scripts

Beyond the per-provider adapters, providers/ ships these utilities:

  • get-local-analytics <codex|opencode> — local harness telemetry (7-day chart, top models, top projects, per-window token breakdown), cached 120s. Shares its windowing with scripts/local-analytics.jq.
  • local-cost-common — sourced helper building the Hermes cost SQL, so the summary and analytics adapters cannot drift on which ledger rows count as a known cost.
  • get-usage-history — prints the local usage history written by the dispatcher (~/.cache/AiOverviewControl/usage-history.jsonl), trimmed by AIOC_HISTORY_MAX.
  • export-usage-history — copies that history to CSV or JSONL (see configuration).
  • get-provider-wrapper — shared single-provider wrapper behind the get-<id>-usage entrypoints.
  • send-quota-alert — deduplicated desktop notification sender used by quota alerts (flock + notify-send).

Direct tests

Every provider is covered by providers/get-provider-health. Most providers also have a normalized providers/get-<id>-usage entrypoint; use the dispatcher for Claude and providers/get-pi-analytics for pi. Examples:

./providers/get-codex-usage | jq .
./providers/get-openrouter-usage | jq .
./providers/get-ollama-usage | jq .
./providers/get-xai-usage | jq .
./providers/get-kimi-usage | jq .
./providers/get-provider-health "codex,openrouter,ollama,xai,kimi" | jq .
./providers/get-provider-usage "gemini,cloudflare,mistral,glm,nvidia,minimax,kimi,qwen,xai,kilo,kiro" | jq .

Output schema

{
  "provider": "codex",
  "source": "codex-app-server",
  "usage": {
    "identity": {
      "providerID": "codex",
      "accountEmail": "account@example.com",
      "loginMethod": "plus"
    },
    "primary": {
      "usedPercent": 25,
      "windowMinutes": 300,
      "resetsAt": "2026-06-11T01:33:15Z",
      "resetDescription": "Session"
    },
    "secondary": null,
    "tertiary": null,
    "updatedAt": "2026-06-10T20:44:55Z"
  },
  "credits": { "remaining": "0" }
}

Errors use {provider, source, error:{code, kind, message}}. A provider without a quota API uses a normal informational usage object (via json_note_usage) rather than an error — primary.resetDescription carries the human label and primary.displayValue (optional) carries a string to render.