PaperScope
LIVE · 2026-09-11 05:40 UTC

Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural Understanding

Michael Picheny

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.11772 v1
Category
Submitted
2026-09-10

Abstract

Cross-cultural understanding has become increasingly important in today's highly connected, cross-national world. The success of LLM-based technologies is now driving the development of automated tools to aid understanding for nonnative people trying to succeed in cross-cultural environments. Building such automated tools is often done by leveraging in-thewild text, audio, and video data. This paper presents techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools. The focus is on processes and speech tools that can easily be used by cross-cultural tool builders without requiring deep speech processing expertise. Using publicly available videos from YouTube and Whisper-based tools, average transcription error rate across seven languages (Spanish, Japanese, Korean, Mandarin, Turkish, Russian, and Hebrew) of 30% are observed. With a modest amount of fine-tuning data, the average error rate can be reduced to 20% making such output much more usable for downstream processing. Speech and metadata associated with these videos that can be used by the community to further refine these experiments are released as well.

Comment: 7 pages, 2 figures, 5 tables

arXiv abs page · PDF · same-day batch