Automated Researchers Can Reliably Mitigate Alignment Failures
Chen Yueh-Han · Jiaxin Wen · Jan Hendrik Kirchner
Claude Opus 4.8 agents ("automated alignment researchers") reliably discover post-training methods that mitigate ten alignment failures — deception, sycophancy, jailbreaks, prompt injection, and more — while preserving general capability. Methods generalize to held-out benchmarks, to multi-turn behavioral audits, and to models up to 4.7× larger than the target model, outperforming ideas from 28 experienced human researchers within an average of 6 hours of search.