PaperScope
LIVE · 2026-09-03 05:40 UTC

Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00330 v1
Category
Submitted
2026-08-27

Abstract

In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speech Recognition) transcripts of spontaneous phone conversations, which can be unclear, repetitive, and mostly lack punctuation. To systematically study this real-world task, we curate a human-annotated topic-utterance judgments dataset sourced from real call-center transcripts. We compare three types of matchers: a regex-based baseline, zero-shot sentence embedding encoders, and Gemini-based LLM matchers. In addition, two types of topic representations are studied in our benchmark:keyphrases and natural language description. Our empirical experiments highlight the superior performance of lightweight LLM matchers over embedding and regex models when equipped with natural language descriptions.

Comment: Accepted at the 11th Workshop on Natural User-generated Text (W-NUT 2026), EMNLP 2026. Camera-ready version. 9 pages, 2 figures, 3 tables

arXiv abs page · PDF · same-day batch