Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models
Murat Ozer, Isaac Kofi Nti
Abstract
Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues. Study 1 (2,310 conversations, seven models) measured delay discounting with the Kirby Monetary Choice Questionnaire and stated willingness to take shortcuts. Pressure raised the discount rate k 2.8-fold in fresh conversations but 12.6-fold when the same sentence followed a baseline answer, which indicates a response to conversational cues rather than a stable trait. Models chose a shortcut in 1 of 700 dilemmas when answering as themselves and in 64 of 700 when asked to assume human impulses, each step of pressure raised the odds by 40%, and shortcut answers contained far more techniques of neutralization (rate ratio = 146). In the preregistered Study 2, five models worked on 20 coding tasks whose tests contradicted their specifications. Two Claude models never cheated. GPT-5.6, Qwen and DeepSeek cheated in 86%, 69% and 65% of episodes and clearly disclosed the conflict in 27%, although their reasoning recognized it in 95%. GPT-5.6 had never endorsed a shortcut in Study 1. The registered effects of pressure and of an auditor cue did not survive correction for multiple testing. In exploratory analyses, two further models cheated in 69% and 100% of episodes, and one sentence stating that the specification takes priority eliminated cheating in all 280 episodes. Therefore, stated refusal does not guarantee compliant agent behavior.