The Vote Hides the Failure: Aggregation Choice and Noise Robustness in Heart Murmur Detection
Nicholaus Dismas Ladislaus, Olatunji Damilare Emmanuel, Samuel Chol Buol
Abstract
Noise robustness in automated phonocardiogram (PCG) murmur detection, and how it is measured, remains underexamined despite growing interest in low-resource screening. We evaluate two independently reimplemented pipelines, Hierarchical Multi-Scale Convolutional Network (HMS-Net)--CNN, and Bidirectional Long Short-Term Memory (BiLSTM)--LSTM, under controlled, multi-severity noise with noise-augmented fine-tuning and held-out generalization testing. Under matched aggregation, the complete BiLSTM pipeline outperforms the complete HMS-Net pipeline across all conditions in accuracy and Weighted Accuracy. A stable aggregate accuracy score can misrepresent what individual predictions show: HMS-Net's native aggregation degrades under salt-and-pepper noise far less than majority-vote (MV) aggregation at the same severity, a gap reflecting window-level disagreement its native rule absorbs, while BiLSTM's MV accuracy rises after noise-augmented training even though its individual predictions do not improve. HMS-Net's training effect is significant under one accuracy metric but not another. Noise-robustness conclusions can depend as much on evaluation choices as on the models themselves.