PaperScope
LIVE · 2026-09-09 05:40 UTC

Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

Kevin Baum, Rūta Binkytė, Felix Jahn

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07627 v1
Category
Submitted
2026-09-07

Abstract

AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.

Comment: 8 pages, currently under review at the NeurIPS 2026 workshop "Foundations of Agentic Systems Theory (FAST)"

arXiv abs page · PDF · same-day batch