PaperScope
LIVE · 2026-09-29 05:40 UTC

Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion

Yasuto Hoshi, Daisuke Miyashita, Jun Deguchi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34063 v1
Category
Submitted
2026-09-28

Abstract

Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the final model output. We study training-free RESA and learned Top-K+$φ$ with frozen backbone language models. A prespecified single-layer screen yields two Qwen3-0.6B/Multi-LexSum interventions for which direct-runtime measurements show positive prespecified request-aggregate local reconstruction gain but worse final KL fidelity than the corresponding all-abstain Exact Top-K baseline on both discovery and prompt-token-disjoint holdout requests. Exact restoration at the same layer instead improves final fidelity, showing that the reversal is specific to approximate completion in these cases. In complementary multi-layer experiments, a task-independent local diagnostic often repairs the tested completion estimators, although the repaired models do not consistently outperform Exact Top-K. Together, these results show that better local reconstruction need not translate into better final-model fidelity.

arXiv abs page · PDF · same-day batch