PaperScope
LIVE · 2026-09-15 05:40 UTC

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Raphaël Troncy, Paolo Papotti, Pietro Michiardi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15561 v1
Category
Submitted
2026-09-14

Abstract

Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.

Comment: 22 pages, 7 figures. Extended version of a paper accepted at EvalLLM 2025 (CORIA-TALN 2025)

arXiv abs page · PDF · same-day batch