PaperScope
LIVE · 2026-09-09 05:40 UTC

What Does an LLM-Agent Leaderboard Rank Actually Compare?

Wei-Jung Huang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07785 v1
Category
Submitted
2026-09-07

Abstract

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.

Comment: Accepted at the 13th International Conference on Data Science and Advanced Analytics (IEEE DSAA'2026)

Journal: The 13th International Conference on Data Science and Advanced Analytics (IEEE DSAA'2026)

arXiv abs page · PDF · same-day batch