PaperScope
LIVE · 2026-09-30 05:40 UTC

When Can Prefixes Compile LoRA? Exact Resource-Capped Tests for Frozen Attention

Joyanta Jyoti Mondal, Ibne Farabi Shihab

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36766 v1
Category
Submitted
2026-09-29

Abstract

Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends on the adapter's target through three conditions. First, observability: at one causal readout, every independent key--value prefix sees the content only through the query, attention partition, and value numerator, so a target that differs on two inputs with equal summaries incurs an error floor at every prefix length; norm caps extend this floor to nearly equal summaries. Second, realizability: at a common query, any prefix reduces exactly to two aggregate variables, and the norm-capped optimum is an attained second-order-cone program, also after a fixed output projection; it places two equal-norm rank-one value updates on opposite sides of compilability. Third, implementation: under affine query exposure, $2r$ signed slots approximate a rank-$r$ value update, but their values grow as $O(ε^{-3/2})$, and the construction passes all 400 tolerance checks in float64 yet only 38 in bfloat16. A first-layer GPT-2 readout with fixed token and position meets the common-query condition without clamping activations; at three such heads, the capped optimum leaves 18.4\% to 74.2\% of the projected adapter effect uncompiled, with a head-dependent value--query ordering. All claims concern local approximation at one head, not whole-network equivalence.

Comment: 26 pages, 4 figures

arXiv abs page · PDF · same-day batch