PaperScope
LIVE · 2026-09-29 05:40 UTC

Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

Etienne Boursier, Nicolas Flammarion

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34667 v1
Category
Submitted
2026-09-28

Abstract

Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.

arXiv abs page · PDF · same-day batch