PaperScope
LIVE · 2026-10-08 05:40 UTC

Lightweight and Versatile Learned Optimization by Recombination of Gradient History

Minyoung Choi, Dalta Imam Maulana, Wanyeong Jung

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09604 v1
Category
Submitted
2026-10-07

Abstract

This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.

arXiv abs page · PDF · same-day batch