PaperScope
LIVE · 2026-09-28 05:40 UTC

Accounting for Bias Enables Sustainable LLM Evaluation

Harshita Katoch, David Antony Selby, Gerrit Großmann, Sebastian Vollmer

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.31184 v1
Submitted
2026-09-25

Abstract

LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.

Comment: 8 pages, 2 figures; SuRE'26: Workshop on Sustainability and Resource-Efficiency of Artificial Intelligence at IJCAI-ECAI 2026

arXiv abs page · PDF · same-day batch