PaperScope
LIVE · 2026-09-29 05:40 UTC

On the Limits of Metacognitive Monitoring in LLMs

Dongqi Han, Yifan Yang, Dongsheng Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34864 v1
Category
Submitted
2026-09-28

Abstract

Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.

arXiv abs page · PDF · same-day batch