How the propose-train-test loop actually works
Anthropic's Automated Alignment Researcher runs a propose-train-test loop: an AI researcher model proposes a mitigation for a specific alignment failure - deception, sycophancy, jailbreaks, reward hacking, privacy violations - trains a checkpoint with that fix, tests it against the target benchmark, then iterates [1]. Across 10 alignment benchmarks, the safety gap closed ranged from 26% to 96%, and the strongest single fix - built from just 2,000 training examples - was roughly 15,000 times more sample-efficient than Anthropic's production alignment procedure [1]. The resulting methods also held up when applied to models roughly 4.7 times larger than the ones they were optimized on, suggesting at least some transfer within the tested benchmark set [1].



