PaperScope
LIVE · 2026-09-03 05:40 UTC

Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

Nishant Balepur, Paiheng Xu, Wei Ai, Eunsol Choi, Rachel Rudinger, Jordan Boyd-Graber

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29887 v1
Category
Submitted
2026-08-30

Abstract

Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.

Comment: EMNLP 2026

arXiv abs page · PDF · same-day batch