# Post-Training

> Post-training is everything done to a model's weights after pretraining: fine-tuning, preference optimization, reinforcement learning and distillation.

Source: https://metavert.io/post-training  
Published: 2026-10-07  
Updated: 2026-10-07

**Post-training** is the collective name for everything done to a language model's weights after large-scale pretraining and before release: supervised fine-tuning on instructions and demonstrations, preference optimization, reinforcement learning against rewards, and distillation. Pretraining produces a next-token predictor with broad knowledge; post-training turns it into an assistant, a reasoner or an agent with particular behaviors.

### The Stages

| Stage | Training signal | What it mainly shapes | Representative source |
| --- | --- | --- | --- |
| Supervised fine-tuning (SFT) | Curated prompt and response pairs | Instruction following, format, tool-call syntax | InstructGPT (Ouyang et al., 2022) |
| Preference optimization | Human or model judgments of which output is better | Helpfulness, tone, refusals | [RLHF](https://metavert.io/rlhf); [DPO](https://metavert.io/direct-preference-optimization) (Rafailov et al., 2023) |
| Reinforcement learning with verifiable rewards | Programmatic checks: a correct answer, passing tests, a grader score | Reasoning, mathematics, code, agentic task completion | DeepSeek-R1 (2025); Tulu 3 (2024) |
| [Distillation](https://metavert.io/knowledge-distillation) | Outputs of a stronger teacher model | Compressing capability into smaller, cheaper models | Hinton et al. (2015) |

The order is not fixed, and labs combine the stages. InstructGPT (March 2022) established the template: supervised learning on labeled demonstrations, then reinforcement learning on human rankings of model outputs. DPO (May 2023) showed that the preference step could be reduced to a simple classification loss with no separate reward model and no sampling during training. The open Tulu 3 recipe (Lambert et al., Allen Institute for AI, November 2024) chains SFT, DPO and a step the authors named reinforcement learning with verifiable rewards.

### The Shift Toward Verifiable Rewards

Preference signals are bounded by what raters can judge. Verifiable rewards are not: if a program can check the answer, the model can practice at a scale no annotation budget allows. DeepSeek-R1 (January 2025) reported that reasoning ability could be "incentivized through pure reinforcement learning" without human-labeled reasoning trajectories, with self-reflection and verification emerging as behaviors, and that the resulting patterns could then be used to improve smaller models. Its training algorithm, [GRPO](https://metavert.io/grpo), came from the same lab's DeepSeekMath paper (February 2024) as a memory-efficient variant of PPO.

The approach has been productized as [reinforcement fine-tuning](https://metavert.io/reinforcement-fine-tuning). OpenAI's guide, as read in October 2026, describes sampling several responses per prompt, scoring them with a programmable grader and applying policy-gradient updates, and lists o4-mini as the only supported model. Its caveats are instructive about the method generally: tasks must be unambiguous and "guess-proof", and a model with a 0% success rate cannot be bootstrapped. Building tasks that meet those conditions is the business of [RL environments](https://metavert.io/rl-environments), and failing to meet them produces [reward hacking](https://metavert.io/reward-hacking). The [supervised versus reinforcement fine-tuning](https://metavert.io/compare/supervised-fine-tuning-vs-reinforcement-fine-tuning) comparison covers when each applies.

### Why Capability Gains Are Attributed to This Stage

The public evidence is suggestive rather than complete. InstructGPT reported that outputs from a 1.3-billion-parameter post-trained model were preferred by human raters over those of the 175-billion-parameter GPT-3, a hundredfold size difference overcome by post-training alone. Tulu 3 applied an open recipe to Llama 3.1 base models and reported results surpassing the official instruct versions built on those same weights, which isolates the recipe as the variable. DeepSeek-R1 attributed its reasoning gains to the reinforcement stage rather than to a new pretraining run. Meta's Llama 3 paper (July 2024) treats pretraining and post-training as distinct phases and released both versions of its largest model.

Two cautions apply. First, the Tulu 3 authors describe post-training data and recipes as "simultaneously the most important pieces of the puzzle and the portion with the least transparency"; frontier labs disclose little, so claims about what share of a commercial model's capability comes from which stage cannot be checked from outside. Second, the stages are not independent. Reinforcement learning amplifies abilities a base model already partly has, as the bootstrapping caveat above implies, so better pretraining and better post-training compound rather than substitute.

### Practical Consequences

Post-training is also where organizations outside the frontier labs can participate. Pretraining a competitive model is out of reach for most; fine-tuning an [open-weight model](https://metavert.io/open-weight-models) with [LoRA](https://metavert.io/lora), a preference dataset or a custom grader is not. The stage carries its own risks: narrow optimization can erode general ability, the forgetting problem studied in [continual learning](https://metavert.io/continual-learning); training on model-generated data raises [model collapse](https://metavert.io/model-collapse) concerns; and every reward is a specification that can be gamed. Evaluation with held-out tasks, which Tulu 3 formalized as separate development and unseen evaluation suites, is the main defense.

## Related Topics

- [Fine-Tuning](https://metavert.io/fine-tuning) — The supervised core of post-training
- [RLHF](https://metavert.io/rlhf) — The original preference-based stage
- [Direct Preference Optimization](https://metavert.io/direct-preference-optimization) — Preference learning without a reward model
- [Reinforcement Fine-Tuning](https://metavert.io/reinforcement-fine-tuning) — Training against graders and verifiable rewards
- [GRPO](https://metavert.io/grpo) — The policy-gradient algorithm behind DeepSeek-R1
- [Knowledge Distillation](https://metavert.io/knowledge-distillation) — Transferring post-trained behavior into smaller models
- [RL Environments](https://metavert.io/rl-environments) — Where verifiable rewards come from
- [Reward Hacking](https://metavert.io/reward-hacking) — The characteristic failure of reward-driven stages
- [AI Model Training](https://metavert.io/ai-model-training) — The full pipeline, including pretraining

## Further Reading

- [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155) — Ouyang et al., arXiv, March 2022
- [Direct Preference Optimization: Your Language Model is Secretly a Reward Model](https://arxiv.org/abs/2305.18290) — Rafailov et al., arXiv, May 2023
- [Tulu 3: Pushing Frontiers in Open Language Model Post-Training](https://arxiv.org/abs/2411.15124) — Lambert et al., arXiv, November 2024
- [DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning](https://arxiv.org/abs/2501.12948) — DeepSeek-AI, arXiv, January 2025
- [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://arxiv.org/abs/2402.03300) — Shao et al., arXiv, February 2024
- [The Llama 3 Herd of Models](https://arxiv.org/abs/2407.21783) — Meta, arXiv, July 2024
- [Reinforcement fine-tuning guide](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning) — OpenAI documentation, accessed October 2026
- [Distilling the Knowledge in a Neural Network](https://arxiv.org/abs/1503.02531) — Hinton, Vinyals and Dean, arXiv, March 2015
