PaperScope
LIVE · 2026-09-21 05:40 UTC

Samsone: A Family of Open Small Audio Language Models for On-Device Inference

Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.21666 v1
Category
Submitted
2026-09-18

Abstract

The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.

Comment: Accepted for Interspeech 2026

arXiv abs page · PDF · same-day batch