Mechanistic Effect of RL Post-Training on Pretrained Policies

Characterize how reinforcement learning with verifiable rewards changes the behavior of pretrained large language models for reasoning by determining whether RL primarily sharpens existing action preferences or discovers new, previously low-probability solution modes, and identify the conditions under which each behavior occurs.

Background

There are competing views on how RL affects pretrained reasoning models: some studies suggest RL mainly sharpens existing preferences, while others argue RL composes new skills. The lack of step-level supervision and massive action spaces in natural language make this mechanistic question difficult to resolve.

A precise characterization of how RL reshapes the inherited policy—at the level of distributions over actions and the structure of reasoning traces—would clarify when RL provides genuine discovery versus amplification and guide post-training strategies.

References

As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model?

Understanding Reasoning from Pretraining to Post-Training  (2607.16097 - Shen et al., 17 Jul 2026) in Abstract