Learning Dynamics in RL Post-Training for Language Models

Title:Learning Dynamics in RL Post-Training for Language Models

Abstract:Reinforcement learning (RL) post-training is a critical stage in modern language model development, playing a key role in improving alignment and reasoning ability. However, several phenomena remain poorly understood, including the reduction in output diversity. To gain a broader understanding of RL post-training, we analyze the learning dynamics of RL post-training from a perspective that has been studied in supervised learning but remains underexplored in RL. We adopt an empirical neural tangent kernel (NTK) framework and decompose the NTK into two components to characterize how RL updates propagat…

Title:Learning Dynamics in RL Post-Training for Language Models

View PDF HTML (experimental)

Abstract:Reinforcement learning (RL) post-training is a critical stage in modern language model development, playing a key role in improving alignment and reasoning ability. However, several phenomena remain poorly understood, including the reduction in output diversity. To gain a broader understanding of RL post-training, we analyze the learning dynamics of RL post-training from a perspective that has been studied in supervised learning but remains underexplored in RL. We adopt an empirical neural tangent kernel (NTK) framework and decompose the NTK into two components to characterize how RL updates propagate across training samples. Our analysis reveals that limited variability in feature representations can cause RL updates to systematically increase model confidence, providing an explanation for the commonly observed reduction in output diversity after RL post-training. Furthermore, we show that effective learning in this regime depends on rapidly shaping the classifier, which directly affects the gradient component of the NTK. Motivated by these insights, we propose classifier-first reinforcement learning (CF-RL), a simple two-stage training strategy that prioritizes classifier updates before standard RL optimization. Experimental results validate our theoretical analysis by demonstrating increased model confidence and accelerated optimization under CF-RL. Additional analysis shows that the mechanism underlying CF-RL differs from that of linear-probing-then-fine-tuning in supervised learning. Overall, our study formalizes the learning dynamics of RL post-training and motivates further analysis and improvement.


Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2601.04670 [cs.LG]
	(or arXiv:2601.04670v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2601.04670 arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Akiyoshi Tomihari [view email] [v1] Thu, 8 Jan 2026 07:32:15 UTC (515 KB)

Title:Learning Dynamics in RL Post-Training for Language Models

Title:Learning Dynamics in RL Post-Training for Language Models

Submission history

Similar Posts