PaperScope
LIVE · 2026-09-29 05:40 UTC

ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs

Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32344 v1
Category
Submitted
2026-09-26

Abstract

For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confidence, and relation metadata; a single ranking supports multiple write budgets while preserving all facts in external memory. On CounterFact with Qwen3-4B, ALLOT reaches 0.760 accuracy at a 20% parametric-write budget and recovers 78.4% of the budget-matched oracle gain, with 80% fewer parametric writes than dual-writing every fact. At this budget, jointly adding retrieval and relation features to text improves normalized oracle gain by 6.2 percentage points. Complementary Qwen3-0.6B shared-store results achieve dual-write-level accuracy with 6-14.5% parametric writes, and cross-benchmark transfer retains approximately 88% of in-domain gain. These results support allocating adaptation capacity according to its incremental value rather than treating every factual update as an equally valuable training target.

arXiv abs page · PDF · same-day batch