PaperScope
LIVE · 2026-09-09 05:40 UTC

The Internal Anatomy of Strategic Choice in Large Language Models

Vinícius Ferraz, Leon Houf, Enrico Ferrea

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07478 v1
Category
Submitted
2026-09-07

Abstract

Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.

arXiv abs page · PDF · same-day batch