PaperScope
LIVE · 2026-10-08 05:40 UTC

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Hanzuo Liu, Chunyu Liu, Chaofan Lin, Alex Lamb, Mingyu Gao

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10058 v1
Category
Submitted
2026-10-07

Abstract

Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

Comment: 17 pages, 3 figures, 7 tables

arXiv abs page · PDF · same-day batch