PaperScope
LIVE · 2026-09-15 05:40 UTC

Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries

Frank Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15021 v1
Category
Submitted
2026-09-14

Abstract

Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas sharing a 256 GiB LMCache pool. After adopting an existing packed-page patch, we isolate a raw-pointer fallback that omits the dependency on the current CUDA stream. Controlled byte tests fail under an imposed delay and pass when the dependency is restored; the existing mixed allocator provides a working deployment path. Full-pool allocation checks and service regression complete the validation. A four-block OFF-ON-ON-OFF comparison contains 768 measured requests within two block pairs. Median cross-replica time to first content token falls from 31.715 to 0.605 seconds at 128k input and from 92.047 to 0.790 seconds at 256k. Six-turn synthetic sessions alternating replicas improve by approximately 35% and 45% at initial contexts of 32k and 128k, while fixed placement shows little benefit. This engineering case study identifies practical validation steps and the locality conditions in which shared caching pays off.

Comment: Technical report. 8 pages, 4 figures, 4 tables

arXiv abs page · PDF · same-day batch