Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
Luca Bompani, Marco Fariselli, Giovanni Oltrecolli, Francesco Conti
Abstract
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15~ms per-frame latency while refreshing a complete set of CNN-derived weights every 564~ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65\%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26~dB, and a wideband PESQ of 3.676 at an average power of 45.9~mW. A speech activity detection (SAD) module (98.5\% accuracy, 0.62~mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16~h all-day target on a 100~mAh battery, with an estimated lifetime of up to $\sim$20~h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.