PaperScope
LIVE · 2026-09-03 05:40 UTC

Asymmetries in Spontaneous and Instructed Deception

Josiah Luikham

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.00180 v1
Category
Submitted
2026-08-31

Abstract

Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.

arXiv abs page · PDF · same-day batch