Claude is Now Alignment Pretrained (opens in new tab)
Anthropic are now actively using the approach to alignment often called “Alignment Pretraining” or “Safety Pretraining” — using Stochastic Gradient Descent on a large body of natural or synthetic documents showing the AI assistant doing the right thing. They tried this out, ound it works well, and are now using it.I’m absolutely delighted. I’ve been advocating this approach on LessWrong and the Alignment Forum for several years:How to Control an LLM's Behavior (why my P(DOOM) went down)Motiva...
Read the original article