PaperScope
LIVE · 2026-09-29 05:40 UTC

SAGE: Semantic Audio Generative Encoder

Francesco Brigante, Luca Cerovaz, Davide Marincione, Giorgio Strano, Luca Zhou, Emanuele Rodolà, Michele Mancusi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32755 v1
Submitted
2026-09-26

Abstract

Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.

Comment: 18 pages, 6 figures, 11 tables. Code and weights: https://github.com/francescobrigante/SAGE. Project page: https://sage-music.pages.dev/

arXiv abs page · PDF · same-day batch