PaperScope
LIVE · 2026-09-29 05:40 UTC

LLMs learn different forms of metacognition when trained to predict their own accuracy

Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33886 v1
Category
Submitted
2026-09-27

Abstract

Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.

arXiv abs page · PDF · same-day batch