PaperScope
LIVE · 2026-09-29 05:40 UTC

CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations

Jae Min Woo, Kyongmin Kong, Bogyung Jeong, Minjeong Kim, HaeJun Yoo, Du-Seong Chang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33433 v1
Submitted
2026-09-27

Abstract

Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by Recall@k. Using Pearson's correlation coefficient r, we find that RankDrop is weakly associated with raw text space movement (r=0.084), but strongly associated with Target Alignment Loss and Target Boundary Margin Degradation (r=0.508 and r=0.615). The same pattern appears in OEA retrievers, where RankDrop is better explained by boundary degradation (r=0.472/0.478) than by query movement (r=0.084/0.046). Overall, these results suggest that robust T2A retrieval requires preserving the target's boundary advantage over competing audio under reformulation.

Comment: Accepted to Findings of IJCNLP-AACL. Code and data: https://github.com/sogang-isds/CORA

arXiv abs page · PDF · same-day batch