PaperScope
LIVE · 2026-09-03 05:40 UTC

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

Mingqi Gao, Anthony Sicilia, Weiyan Shi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.26638 v2
Category
Submitted
2026-08-27

Abstract

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.

Comment: Accepted at EMNLP 2026 (Main). Code available: https://github.com/CHATS-lab/ppi-eval

arXiv abs page · PDF · same-day batch