PaperScope
LIVE · 2026-09-29 05:40 UTC

Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports

Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32449 v1
Category
Submitted
2026-09-26

Abstract

Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.

arXiv abs page · PDF · same-day batch