Dynamic Routing as a New Dimension for Test-time Versatility of LLMs
Michal Štefánik, Marek Kadlčík, Josef Kuchař, Michal Spiegel
Abstract
Beyond scaling their parameters and data, large language models currently gain versatility on new problems along a single axis: the tokens they spend on chain-of-thought (CoT). We investigate whether dynamic routing programs, which execute a subset of the model's layers or iterate some of them, can open a second axis of test-time adaptation, complementary to CoT and free of any gradient update. Prior work showed that such programs exist and bring accuracy and efficiency gains on problems similar to those they were trained on; we ask whether they can also be identified rapidly, from a handful of demonstrations (3 or 10), by a strategy that transfers across models and tasks without training. First, we find that strategies that select programs by the probability they assign to the demonstrations' labels, arbitrated by the model's own confidence, bring consistent gains: on average over the 49 tasks of MMLU and substantially on four of seven models, and most of all on far out-of-distribution tasks such as ARC-AGI, where programs double the accuracy of a 7B model whose CoT fails. Second, on MMLU across the seven post-trained models, routing complements CoT in practice: the two succeed on different queries, and their composition exceeds CoT alone. Despite these gains, our analyses show that confidence-based selection leaves much of the potential untapped, in two places in particular: (1) in surfacing the routing potential that is already present early in pre-training but becomes harder to select after post-training, and (2) in making models robust to the refinements routing introduces, since unsuccessful routes tend to drive the residual stream out of the distribution that the following layers expect. Together, our results point to dynamic routing as a paradigm for extending the plasticity of existing and future LLMs in rapid test-time adaptation.