KV Cache Quantization
November 5, 2025 · View on GitHub
-
Quantize What Counts: More For Keys, Less For Values
- TL;DR: Keys carry more information than values → allocate more bits to keys than values in KV cache quantization.
-
Norm-Aware KVQuant: Precision Where It Counts
- TL;DR: Performs layer-wise bit allocation for KV cache quantization, guided by the norm and spectral gap between keys and values.
-
- Please use the
kvqbranch for evaluation. - KVQ argument for lm_eval:
--kvq '{"budget":4,"bit_range":[1,1.58,2,4,6,8],"axis":{"k":1,"v":1},"residual_length":{"k":64,"v":64},"group_budget":64,"group_range":[32, 64, 128]'}
- Please use the
Quantize What Counts
Key-Value Norm Disparity
The expected Frobenius norm of key projections is consistently greater than that of value projections.
$\mathbb{E}[\|W_K\|_F] \;>\; \mathbb{E}[\|W_V\|_F]$
This indicates that key matrices carry richer and more substantial information compared to value matrices.
Corollary
$\mathbb{E}[\|K\|_F] \;>\; \mathbb{E}[\|V\|_F]$
Key-Driven Quantization
Allocating higher bit precision to keys rather than values significantly improves quantization efficiency. Formally, for bit allocations :
If , then the expected inference accuracy under allocation is strictly higher than that under the reversed allocation .
Norm-Aware KVQuant
Supported backends
kvq supports the following models:
-
Qwen 3.0
- Qwen/Qwen3-0.6B
- Qwen/Qwen3-4B
- Qwen/Qwen3-8B
- Qwen/Qwen3-14B
- Qwen/Qwen3-32B
-
Llama 3.0
- meta-llama/Meta-Llama-3-8B
- meta-llama/Meta-Llama-3-8B-Instruct
-
Llama 3.2
- meta-llama/Llama-3.2-1B
- meta-llama/Llama-3.2-1B-Instruct
- meta-llama/Llama-3.2-3B
- meta-llama/Llama-3.2-3B-Instruct
-
Llama 3.3
- meta-llama/Llama-3.3-70B-Instruct
KVQ can be installed via pip:
pip install kvq
Usage
import torch
from kvq import KVQConfig, KVQ
config = KVQConfig(
budget = 4,
model="meta-llama/Llama-3.1-8B-Instruct",
residual_length=32,
group_size={"k": 64, "v": 64}, # Group size for keys and values
axis={"k": 0, "v": 0}, # Axis along which to quantize
)
kv_cache = KVQ(config)
text = "What is the meaning of life?"
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
past_key_values=kv_cache,
use_cache=True,
pad_token_id=tokenizer.eos_token_id,
)
Citation
If you find our method useful, please kindly cite our paper.
@article{hariri2025quantize,
title={Quantize What Counts: More for Keys, Less for Values},
author={Mohsen Hariri and Alan Luo and Weicong Chen and Shaochen Zhong and Tianyi Zhang and Qifan Wang and Xia Hu and Xiaotian Han and Vipin Chaudhary},
year={2025},
journal={arXiv preprint arXiv:2502.15075},
eprint={2502.15075},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.15075},
}
License
The code is released under the MIT License.