PaperScope
LIVE · 2026-10-06 05:40 UTC

AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

Steven Moore, Nicholas Diana

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04267 v1
Category
Submitted
2026-10-03

Abstract

Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.

Comment: 8 pages, 2 tables, Full paper accepted to the AIME Conference 2026

arXiv abs page · PDF · same-day batch