overview.mdx
August 27, 2026 ยท View on GitHub
KV cache offloading moves reusable KV blocks from GPU memory to host memory or storage. This tutorial starts an aggregated Dynamo deployment with one offloading backend, sends requests through it, and checks that the backend is running.
Complete the local installation and start Dynamo's NATS and etcd services before continuing. Run the commands from the root of the cloned Dynamo repository.
If docker run opened a shell in a prebuilt runtime container, you are already in the correct
environment and can skip package installation when the table below marks the backend as included.
For a container running in the background, open a shell with docker exec; the repository is
available at /workspace:
docker exec -it <container-name> bash
cd /workspace
If you installed Dynamo into a Python environment, activate that environment before continuing.
Select one offloading backend for each worker. LMCache, FlexKV, and HiCache are alternatives, not cache layers to stack together. KV-aware routing and NIXL-based prefill/decode transfer are separate features that can operate alongside the selected backend.
| Backend | Engine | Cache tiers | Prebuilt Dynamo runtime container | Python environment |
|---|---|---|---|---|
| LMCache | vLLM | Host and persistent L2 adapters | Inherited from the upstream vLLM image; verify compatibility for the image tag and platform | Install lmcache |
| HiCache | SGLang | Host and optional external storage | Included with SGLang | Included with ai-dynamo[sglang] |
| FlexKV | vLLM in this tutorial | Host, SSD, cloud storage | Requires a custom image or a source build inside the container | Build FlexKV from source |
Choose a backend and follow its tab. The primary examples use Qwen/Qwen3-0.6B and one GPU; the
optional disaggregated FlexKV example uses two GPUs. The vLLM launch scripts start the Dynamo
frontend and workers, and stop all processes when you press Ctrl+C.
The current Dynamo vLLM runtime is based on an upstream vLLM image that provides LMCache on supported platforms. Confirm that the selected image contains a compatible build:
python -c "import lmcache; print(lmcache.__version__)"
For a compatible x86_64 Python environment, install LMCache with:
uv pip install lmcache
LMCache publishes x86_64 wheels built for specific CUDA versions. For Arm64 or a mismatched PyTorch/CUDA stack, build LMCache from source by following the LMCache installation guide.
Start the out-of-process LMCache server and an aggregated vLLM worker:
cd examples/backends/vllm
LMCACHE_L1_SIZE_GB=16 ./launch/agg_lmcache_mp.sh
The script starts lmcache server, waits for its health endpoint, and connects the vLLM worker with
LMCacheMPConnector. Inspect its metrics from another terminal:
curl -s localhost:8080/metrics | grep '^lmcache_mp_'
For server flags and persistent L2 adapters, see the LMCache MP configuration reference.
HiCache is part of SGLang, so both the SGLang runtime container and a Python environment installed
with ai-dynamo[sglang] require no additional package.
In one terminal, start an SGLang worker with a host cache twice the size of its GPU cache:
python -m dynamo.sglang \
--model-path Qwen/Qwen3-0.6B \
--page-size 64 \
--enable-hierarchical-cache \
--hicache-ratio 2 \
--hicache-write-policy write_through \
--hicache-storage-backend nixl \
--skip-tokenizer-init
In another terminal in the same environment, start the frontend:
python -m dynamo.frontend --http-port 8000
Press Ctrl+C in both terminals when you finish.
These flags belong to SGLang and pass through Dynamo unchanged. For external storage and tier-aware shared-cache routing, see SGLang HiCache. That advanced path has additional SGLang-version and Mooncake requirements.
Experimental. The stock Dynamo vLLM runtime does not explicitly install FlexKV. Build it in the active Python environment or create a derived runtime image containing the build:
git clone https://github.com/taco-project/FlexKV.git
cd FlexKV
./build.sh
cd ..
For SSD offloading, also install liburing-dev and libxxhash-dev in the host or image before
building FlexKV.
Start an aggregated vLLM worker with a 32 GB host cache:
cd examples/backends/vllm
FLEXKV_CPU_CACHE_GB=32 ./launch/agg_flexkv.sh
The script sets DYNAMO_USE_FLEXKV=1 and selects FlexKVConnectorV1.
To test the experimental disaggregated FlexKV topology on two GPUs, run:
cd examples/backends/vllm
./launch/disagg_flexkv.sh
The script starts the decode worker with NixlConnector and the prefill worker with PdConnector,
which combines FlexKVConnectorV1 for offloading with NixlConnector for prefill/decode transfer.
It selects the prefill role with --disaggregation-mode prefill. The --is-prefill-worker flag
it replaced was removed in v1.4.0.
For SSD and distributed-cache configuration, see the FlexKV documentation.
After the frontend listens on port 8000, send a request from another terminal:
curl localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Explain how a KV cache reduces repeated computation."}],
"stream": false,
"max_tokens": 30
}'
Send the request again to create an opportunity to reuse its prompt prefix. A short request confirms the deployment path but does not establish a performance improvement. Use a representative workload with long, repeated prefixes to measure Time To First Token (TTFT) and cache hit rate.
Next Steps
- Add KV-aware routing when multiple workers can reuse cached prefixes.
- Combine one offloading backend with disaggregated serving. In that topology, the offloading connector manages storage while NIXL transfers blocks from prefill to decode workers.
- To make a custom inference engine visible to the KV router, implement KV event publishing.