PaperScope
LIVE · 2026-09-21 05:40 UTC

The Spoken Wikipedia Presentation Corpus

Thomas Ranzenberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.21676 v1
Submitted
2026-09-18

Abstract

We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.

Comment: Accepted at SLT 2026

arXiv abs page · PDF · same-day batch