« All posts

Aligning SFT with RL through MCMC Sampling of Training Data

Discover how MCMC sampling can align SFT with RL, enhancing model performance and generalization.

Enhancing frontier models via post-training typically involves supervised finetuning (SFT) and reinforcement learning (RL). While RL is known for its strong generalization capabilities, SFT often suffers from catastrophic forgetting. This research introduces a Markov chain Monte Carlo (MCMC) sampling algorithm that adjusts off-policy data to better support SFT, enabling it to compete with leading post-training methods by improving generalization and reducing forgetting.

This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work