Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines (opens in new tab)
Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime is not always classical catastrophic forgetting: a pipeline may preserve previously learned behaviors while still failing to accumulate reusable methodological knowledge about how to train the next campaign. We call this failure mode scientific amnesia. This paper tur...
Read the original article