PaperScope
LIVE · 2026-09-29 05:40 UTC

Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device

Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35005 v1
Category
Submitted
2026-09-28

Abstract

Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.

Comment: Interspeech 2026, 5 pages, 2 figures

Journal: Proc. Interspeech 2026, 4077-4081

arXiv abs page · PDF · same-day batch