PaperScope
LIVE · 2026-09-15 05:40 UTC

A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset

Qiyang Sun, Langqing Zhang, Yupei Li, Björn Schuller

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14145 v1
Submitted
2026-09-12

Abstract

Kinship verification is a task involving determining whether two individuals share a first-order kin relation. To tackle this task, we propose CONVTRAP-TN, a new architecture for audio-based kinship verification, and conduct an ablation study on the proposed model. To the best of our knowledge, we are the first to apply the successful transformer architecture to the task of audio-based kinship verification. Furthermore, we also collect a custom speech dataset, ARKIN, which accurately reflects everyday recording conditions. We do this because only a few speech datasets with kinship labels currently exist, all of which either source extremely noisy in-the-wild data from the internet, or instruct speakers to record in specific environments. These settings fail to reflect real-world scenarios where users record on personal devices under unrestrained conditions. Additionally, we perform a series of preliminary baseline experiments on the collected dataset, including speaker verification and recognition, speech recognition, age estimation, and kinship verification, as well as cross-dataset kinship verification experiments to show that existing methods are not robust across datasets.

Comment: 7 pages, 4 figures. Accepted to IEEE Spoken Language Technology Workshop 2026

arXiv abs page · PDF · same-day batch