A small amount of this data produced broad gains beyond the training scenarios. (opens in new tab)
A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental health. These evals varied widely in domain, task format, and grading scheme.
Read the original article