PaperScope
LIVE · 2026-09-07 05:40 UTC

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

Esteban Gutiérrez, Lonce Wyse, Frederic Font, Xavier Serra

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.04634 v1
Category
Submitted
2026-09-04

Abstract

As generative audio models grow in complexity, the computational and ecological costs of synthesizing everyday sounds have become increasingly prohibitive, often requiring industrial-scale resources and massive datasets. In this paper, we present SCAPES: a Semantically Conditioned Autoregressive Prior for Environmental Sounds. SCAPES is a lightweight, resource-efficient generative model designed to synthesize high-fidelity environmental textures through high-level semantic control. By operating on the continuous latent manifold of a neural audio codec, our approach bypasses the rigid structural constraints inherent to discrete tokenization. We propose a segmentation strategy that decomposes audio into overlapping segments, enabling a Continuous Normalizing Flow (CNF) to model the evolution of latent trajectories using Flow Matching. Our experiments demonstrate that a 36-million parameter instance of SCAPES can be trained on limited, uncurated datasets using a single consumer-grade GPU. Notably, convergence is achieved after training for approximately twice the source audio duration, yielding high-fidelity outputs with robust long-term stability and semantic consistency. Furthermore, we showcase the model's capacity for smooth semantic interpolation, providing a flexible and accessible tool for open research and creative sound design. Code, pretrained weights, audio examples, and an interactive demo are publicly available on our project page https://cordutie.github.io/projects/scapes.html

Comment: Accepted to the Digital Audio Fx (DAFx) Conference 2026 to be held in Cambridge, USA. 8 pages, 3 figures and 2 tables

arXiv abs page · PDF · same-day batch