PaperScope
LIVE · 2026-09-03 05:40 UTC

Automated Researchers Can Mitigate Well-characterized Alignment Failures

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.28945 v3
Category
Submitted
2026-08-28

Abstract

Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers. These results suggest that automating alignment research on well-characterized failures may be practical in the near term.

arXiv abs page · PDF · same-day batch