PaperScope
LIVE · 2026-09-29 05:40 UTC

RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

Ruoling Qi, Yirui Liu, Xuaner Wu, Yuxin Jin, Jian Chen, Jiayu Qin, Yin Chen, Jiawei Shao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33503 v1
Category
Submitted
2026-09-27

Abstract

Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.

arXiv abs page · PDF · same-day batch