PaperScope
LIVE · 2026-09-22 05:40 UTC

Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23152 v1
Category
Submitted
2026-09-19

Abstract

Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.

arXiv abs page · PDF · same-day batch