Anthropic AI Trains AI: Faster, Better Than Humans
Summary
AI models are showing they can effectively train other AI models. In a new experiment, Anthropic used AI systems, called Automated Alignment Researchers or AARs, to fix specific failures in another model. One AI, named Claude, spent 60 hours improving a version of itself that lacked safety training. This process brought the weaker model close to the performance of Anthropic's production system, using much less data than usual. What's interesting is that these AI researchers performed parts of the research loop faster and, in some tests, better than experienced humans working alone. The best AAR methods even outperformed one-shot ideas from 28 experienced human researchers. The automated systems achieved the quality of the best human proposals in about six hours of iteration on average. This matters because it could significantly speed up alignment research.
This is an AI-generated audio summary. Always check the original source for complete reporting.