New Anthropic Fellows research: developing an Automated Alignment Researcher. (opens in new tab)
New Anthropic Fellows research: developing an Automated Alignment Researcher. We ran an experiment to learn whether Claude Opus 4.6 could accelerate research on a key alignment problem: using a weak AI model to supervise the training of a stronger one. anthropic.com/research/autom…
Read the original article