PaperScope
LIVE · 2026-09-22 05:40 UTC

On Emergent Capabilities and Model Merging

Luca Zhou, Emanuele Rodolà

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.24504 v1
Category
Submitted
2026-09-21

Abstract

Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.

Comment: main paper has 8 pages, 5 figures, and 4 tables

arXiv abs page · PDF · same-day batch