Learning to Evict from Key-Value Cache
July 31, 2026 ยท View on GitHub
This software project accompanies the research paper: Learning to Evict from Key-Value Cache by Luca Moschella, Laura Manduchi, Ozan Sener.

We introduce KV Policy (KVP), a framework of lightweight per-head RL agents that learn to rank KV cache entries by their predicted future utility, enabling adaptive eviction without modifying the underlying LLM or adding inference overhead.
Getting Started
Requires Python >= 3.10 and uv.
uv sync
cp .env.template .env # edit with your paths
Download model weights:
uv run tune download Qwen/Qwen2.5-7B-Instruct \
--output-dir models/Qwen2.5-7B-Instruct
Dataset Generation
Extract Q, K, V activations from the RULER dataset:
uv run tune run src/kvcompression/entrypoints/datagen/unroll_and_store.py \
--config configs/preprocess_ruler.yaml \
device=cuda:0 total_num_chunks=1 current_chunk=0
For parallel processing across multiple GPUs, follow the commented-out chunking pattern in configs/preprocess_ruler.yaml.
Copy the pre-generated train/val splits to the output directory:
cp assets/ruler_split/*.txt \
data/gqa_safetensors/simonjegou_ruler/Qwen2.5-7B-Instruct/temperature_0.00/
Training
Qwen2.5-7B-Instruct uses Grouped Query Attention with 28 layers and 4 KV heads, requiring 112 agents (one per layer-head pair).
Train a single agent:
uv run tune run src/kvcompression/entrypoints/rl/train_agent_sampler_distributed.py \
--config configs/train_single_head.yaml distributed=False
Override layer and head:
uv run tune run src/kvcompression/entrypoints/rl/train_agent_sampler_distributed.py \
--config configs/train_single_head.yaml \
target_layer_idx=15 kv_head_idx=2 distributed=False
Each agent is trained with DDP across 8 GPUs. To train all 112 agents, follow the commented-out sweep pattern in configs/train_single_head.yaml.
Trained agents are saved to agents/grouped/<sweep_name>/layer_<NNNNNN>/kv_head_<NNN>/.
Inference
After training, generate a config composite for the trained agents:
uv run python scripts/orchestration/agent_registry.py list-sweeps
uv run python scripts/orchestration/agent_registry.py discover <sweep_name>
uv run python scripts/orchestration/assignment_generator.py write-composites <sweep_name>
See notebooks/inference_demo.ipynb for a demo of generation with learned KV cache compression.
Citation
If you find our work useful, please cite the following paper:
@inproceedings{
moschella2026learningtoevict,
title={Learning to Evict from Key-Value Cache},
author={Luca Moschella and Laura Manduchi and Ozan Sener},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=0OevIlRMYN}
}
Acknowledgements
Our codebase is built using multiple open-source contributions, please see ACKNOWLEDGEMENTS for more details.
License
Please check out the repository LICENSE before using the provided code.