Anthropic paper: automated researchers improved alignment on all 10 benchmarks
A new Anthropic paper reports that automated systems improved a model's scores on ten alignment benchmarks, at roughly $4 per hour versus $150 for human researchers.
Anthropic published a paper on Friday titled "Automated Researchers Can Reliably Mitigate Alignment Failures," describing how AI systems could improve a model's performance on alignment benchmarks. Led by Anthropic fellow Chen Yueh-Han, the work reports that when given 10 benchmarks targeting specific misaligned behaviours, the automated systems improved performance on every one of them without degrading overall performance.
How the system works
According to the paper, the automated system replicates much of the traditional research workflow: it searches the available literature, proposes a method, and trains the model with that method for 30 minutes, gradually raising the benchmark score across several iterations. Effective methods are kept and ineffective ones discarded, letting the system run quickly and at scale.
The paper compares the Automated Alignment Researcher (AAR) directly with its human counterpart. "The best AAR method beats what experienced humans propose, on average within six hours," it states, adding that "human guided research directions do not lead to stronger performance." It also includes a cost comparison: an AAR costs roughly $4 per hour in API inference, against the $150 per hour the company pays its human researchers.
Stated limitations
The authors list constraints as well. The automated system only works insofar as the benchmarks actually reflect the intended alignment goals, and building and maintaining those benchmarks remains substantial work; as does maintaining and expanding the literature the automated researchers draw on. The paper frames the findings as "early evidence that automated alignment post-training could become practical in the near term."
Full details are in the TechCrunch report.