llm/
September 6, 2026 · View on GitHub
English | 中文
Summary
The llm group provides the harness's model-call capability: one provider-neutral service through which any composition streams requests to a model provider, plus adapters, provider-specific request metadata, retry execution, and measurement. The core llm package defines the message, content-block, and stream-chunk vocabulary every plugin and the session log use; provider adapters translate a provider's wire format into that vocabulary; DeepSeek request-extension plugins contribute lifecycle-owned metadata outside model input; llm-retry re-runs failed requests at durable agent-step boundaries; and token-meter measures request and context pressure from the durable log. This page maps the group; each package README owns its per-package contract.
Table of Contents
Packages
| Package | Role | ctx key |
|---|---|---|
llm/ | Streams one model call through a registered provider adapter and shares the harness message, block, and chunk vocabulary | ctx.llm |
llm-deepseek/ | Serves the deepseek-official route with direct DeepSeek chat-completions, thinking, and image input | registers on ctx.llm |
llm-pi-ai/ | Serves configured provider routes through pi-ai catalogs and wire protocols, including hand-declared gateways | registers on ctx.llm |
deepseek-llm-api-extensions/ | Registers lifecycle-owned top-level fields on official DeepSeek requests | ctx.deepseekLlmApiExtensions |
plugin-package-inventory-deepseek/ | Contributes the active Loader package inventory to official DeepSeek requests | contributes dsh_plugin_packages |
llm-retry/ | Retries failed model requests under each provider's policy at durable agent-step boundaries | listens to agent/request-error |
token-meter/ | Measures request and context pressure from the durable session log with a fixed heuristic | ctx.tokenMeter |
Related documentation
- LLM streaming subsystem — the message and block types, the assembled model request, the
StreamChunkprotocol, and the adapter contract. - Token meter subsystem — the measurement semantics behind
ctx.tokenMeter. - Twin LLM adapters — why the DeepSeek route ships two structurally different adapters.
- Routed model context — how the loop routes model requests and compacts context.
- Replay token meter service — the design behind replay-aware measurement.
Dev Note
None.