Post-Training
Post-training is the collective name for everything done to a language model's weights after large-scale pretraining and before release: supervised fine-tuning on instructions and demonstrations, preference optimization, reinforcement learning against rewards, and distillation. Pretraining produces a next-token predictor with broad knowledge; post-training turns it into an assistant, a reasoner or an agent with particular behaviors.
The Stages
| Stage | Training signal | What it mainly shapes | Representative source |
|---|---|---|---|
| Supervised fine-tuning (SFT) | Curated prompt and response pairs | Instruction following, format, tool-call syntax | InstructGPT (Ouyang et al., 2022) |
| Preference optimization | Human or model judgments of which output is better | Helpfulness, tone, refusals | RLHF; DPO (Rafailov et al., 2023) |
| Reinforcement learning with verifiable rewards | Programmatic checks: a correct answer, passing tests, a grader score | Reasoning, mathematics, code, agentic task completion | DeepSeek-R1 (2025); Tulu 3 (2024) |
| Distillation | Outputs of a stronger teacher model | Compressing capability into smaller, cheaper models | Hinton et al. (2015) |
The order is not fixed, and labs combine the stages. InstructGPT (March 2022) established the template: supervised learning on labeled demonstrations, then reinforcement learning on human rankings of model outputs. DPO (May 2023) showed that the preference step could be reduced to a simple classification loss with no separate reward model and no sampling during training. The open Tulu 3 recipe (Lambert et al., Allen Institute for AI, November 2024) chains SFT, DPO and a step the authors named reinforcement learning with verifiable rewards.
The Shift Toward Verifiable Rewards
Preference signals are bounded by what raters can judge. Verifiable rewards are not: if a program can check the answer, the model can practice at a scale no annotation budget allows. DeepSeek-R1 (January 2025) reported that reasoning ability could be "incentivized through pure reinforcement learning" without human-labeled reasoning trajectories, with self-reflection and verification emerging as behaviors, and that the resulting patterns could then be used to improve smaller models. Its training algorithm, GRPO, came from the same lab's DeepSeekMath paper (February 2024) as a memory-efficient variant of PPO.
The approach has been productized as reinforcement fine-tuning. OpenAI's guide, as read in October 2026, describes sampling several responses per prompt, scoring them with a programmable grader and applying policy-gradient updates, and lists o4-mini as the only supported model. Its caveats are instructive about the method generally: tasks must be unambiguous and "guess-proof", and a model with a 0% success rate cannot be bootstrapped. Building tasks that meet those conditions is the business of RL environments, and failing to meet them produces reward hacking. The supervised versus reinforcement fine-tuning comparison covers when each applies.
Why Capability Gains Are Attributed to This Stage
The public evidence is suggestive rather than complete. InstructGPT reported that outputs from a 1.3-billion-parameter post-trained model were preferred by human raters over those of the 175-billion-parameter GPT-3, a hundredfold size difference overcome by post-training alone. Tulu 3 applied an open recipe to Llama 3.1 base models and reported results surpassing the official instruct versions built on those same weights, which isolates the recipe as the variable. DeepSeek-R1 attributed its reasoning gains to the reinforcement stage rather than to a new pretraining run. Meta's Llama 3 paper (July 2024) treats pretraining and post-training as distinct phases and released both versions of its largest model.
Two cautions apply. First, the Tulu 3 authors describe post-training data and recipes as "simultaneously the most important pieces of the puzzle and the portion with the least transparency"; frontier labs disclose little, so claims about what share of a commercial model's capability comes from which stage cannot be checked from outside. Second, the stages are not independent. Reinforcement learning amplifies abilities a base model already partly has, as the bootstrapping caveat above implies, so better pretraining and better post-training compound rather than substitute.
Practical Consequences
Post-training is also where organizations outside the frontier labs can participate. Pretraining a competitive model is out of reach for most; fine-tuning an open-weight model with LoRA, a preference dataset or a custom grader is not. The stage carries its own risks: narrow optimization can erode general ability, the forgetting problem studied in continual learning; training on model-generated data raises model collapse concerns; and every reward is a specification that can be gamed. Evaluation with held-out tasks, which Tulu 3 formalized as separate development and unseen evaluation suites, is the main defense.
Further Reading
- Training language models to follow instructions with human feedback — Ouyang et al., arXiv, March 2022
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafailov et al., arXiv, May 2023
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Lambert et al., arXiv, November 2024
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI, arXiv, January 2025
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Shao et al., arXiv, February 2024
- The Llama 3 Herd of Models — Meta, arXiv, July 2024
- Reinforcement fine-tuning guide — OpenAI documentation, accessed October 2026
- Distilling the Knowledge in a Neural Network — Hinton, Vinyals and Dean, arXiv, March 2015