Aligning SFT with RL through MCMC Sampling of Training Data
Discover how MCMC sampling can align SFT with RL, enhancing model performance and generalization.
Enhancing frontier models via post-training typically involves supervised finetuning (SFT) and reinforcement learning (RL). While RL is known for its strong generalization capabilities, SFT often suffers from catastrophic forgetting. This research introduces a Markov chain Monte Carlo (MCMC) sampling algorithm that adjusts off-policy data to better support SFT, enabling it to compete with leading post-training methods by improving generalization and reducing forgetting.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work