Post-Training

Post-training is the collective name for everything done to a language model's weights after large-scale pretraining and before release: supervised fine-tuning on instructions and demonstrations, preference optimization, reinforcement learning against rewards, and distillation. Pretraining produces a next-token predictor with broad knowledge; post-training turns it into an assistant, a reasoner or an agent with particular behaviors.

The Stages

StageTraining signalWhat it mainly shapesRepresentative source
Supervised fine-tuning (SFT)Curated prompt and response pairsInstruction following, format, tool-call syntaxInstructGPT (Ouyang et al., 2022)
Preference optimizationHuman or model judgments of which output is betterHelpfulness, tone, refusalsRLHF; DPO (Rafailov et al., 2023)
Reinforcement learning with verifiable rewardsProgrammatic checks: a correct answer, passing tests, a grader scoreReasoning, mathematics, code, agentic task completionDeepSeek-R1 (2025); Tulu 3 (2024)
DistillationOutputs of a stronger teacher modelCompressing capability into smaller, cheaper modelsHinton et al. (2015)

The order is not fixed, and labs combine the stages. InstructGPT (March 2022) established the template: supervised learning on labeled demonstrations, then reinforcement learning on human rankings of model outputs. DPO (May 2023) showed that the preference step could be reduced to a simple classification loss with no separate reward model and no sampling during training. The open Tulu 3 recipe (Lambert et al., Allen Institute for AI, November 2024) chains SFT, DPO and a step the authors named reinforcement learning with verifiable rewards.

The Shift Toward Verifiable Rewards

Preference signals are bounded by what raters can judge. Verifiable rewards are not: if a program can check the answer, the model can practice at a scale no annotation budget allows. DeepSeek-R1 (January 2025) reported that reasoning ability could be "incentivized through pure reinforcement learning" without human-labeled reasoning trajectories, with self-reflection and verification emerging as behaviors, and that the resulting patterns could then be used to improve smaller models. Its training algorithm, GRPO, came from the same lab's DeepSeekMath paper (February 2024) as a memory-efficient variant of PPO.

The approach has been productized as reinforcement fine-tuning. OpenAI's guide, as read in October 2026, describes sampling several responses per prompt, scoring them with a programmable grader and applying policy-gradient updates, and lists o4-mini as the only supported model. Its caveats are instructive about the method generally: tasks must be unambiguous and "guess-proof", and a model with a 0% success rate cannot be bootstrapped. Building tasks that meet those conditions is the business of RL environments, and failing to meet them produces reward hacking. The supervised versus reinforcement fine-tuning comparison covers when each applies.

Why Capability Gains Are Attributed to This Stage

The public evidence is suggestive rather than complete. InstructGPT reported that outputs from a 1.3-billion-parameter post-trained model were preferred by human raters over those of the 175-billion-parameter GPT-3, a hundredfold size difference overcome by post-training alone. Tulu 3 applied an open recipe to Llama 3.1 base models and reported results surpassing the official instruct versions built on those same weights, which isolates the recipe as the variable. DeepSeek-R1 attributed its reasoning gains to the reinforcement stage rather than to a new pretraining run. Meta's Llama 3 paper (July 2024) treats pretraining and post-training as distinct phases and released both versions of its largest model.

Two cautions apply. First, the Tulu 3 authors describe post-training data and recipes as "simultaneously the most important pieces of the puzzle and the portion with the least transparency"; frontier labs disclose little, so claims about what share of a commercial model's capability comes from which stage cannot be checked from outside. Second, the stages are not independent. Reinforcement learning amplifies abilities a base model already partly has, as the bootstrapping caveat above implies, so better pretraining and better post-training compound rather than substitute.

Practical Consequences

Post-training is also where organizations outside the frontier labs can participate. Pretraining a competitive model is out of reach for most; fine-tuning an open-weight model with LoRA, a preference dataset or a custom grader is not. The stage carries its own risks: narrow optimization can erode general ability, the forgetting problem studied in continual learning; training on model-generated data raises model collapse concerns; and every reward is a specification that can be gamed. Evaluation with held-out tasks, which Tulu 3 formalized as separate development and unseen evaluation suites, is the main defense.

Further Reading