Supervised Fine-Tuning vs Reinforcement Fine-Tuning
ComparisonSupervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) are the two main ways to change a pretrained model's behaviour during post-training. SFT, the default meaning of fine-tuning, shows the model example outputs and trains it to reproduce them. RFT lets the model produce its own outputs, scores them with a grader, and shifts the weights toward higher-scoring behaviour, typically with a policy-gradient method such as GRPO.
The practical rule is short. Use SFT when good answers can be written down. Use RFT when good answers can be recognised and scored but not easily demonstrated, and when the model can already succeed some of the time. The strongest published systems use both, in sequence.
Feature Comparison
| Dimension | Supervised fine-tuning | Reinforcement fine-tuning |
|---|---|---|
| Training signal | Example outputs: the model is trained to maximise the likelihood of demonstrated responses | A numeric reward from a grader applied to the model's own sampled outputs |
| What must be supplied | Prompt–response pairs | Prompts plus a grader (programmatic check, rubric or model-based scorer) |
| Typical algorithm | Cross-entropy on target tokens | Policy gradient; GRPO, a PPO variant introduced in DeepSeekMath (February 2024), is widely used |
| Data volume (OpenAI guidance) | Minimum 10 examples; improvements typically seen from 50–100 | Start with several dozen to a few hundred; up to 50,000 training examples |
| Performance ceiling | Tied to the quality of the demonstrations | Can exceed the demonstrations: DeepSeek-R1 reports surpassing counterparts trained by supervised learning on human demonstrations for verifiable tasks |
| Starting requirement | None beyond a base model | The model must sometimes succeed: OpenAI advises against RFT when scores sit at the minimum or maximum |
| Characteristic failure | Reproduces the limits and errors of the examples | Reward hacking: scoring well on the grader without being correct (noted in OpenAI's documentation and by DeepSeek) |
| Output style | Follows the demonstrations | Can drift: DeepSeek-R1-Zero showed poor readability and language mixing |
| Small-model evidence | Distillation by SFT: DeepSeek-R1-Distill-Qwen-32B scored 72.6% on AIME 2024 | RL applied directly at the same size: DeepSeek-R1-Zero-Qwen-32B scored 47.0% |
| Compute per example | One forward and backward pass per example | Several sampled responses per prompt, each scored; exact cost ratios are not publicly documented |
| Hosted availability (OpenAI, October 2026) | GPT-4.1, GPT-4.1 mini and GPT-4.1 nano | o4-mini only; OpenAI's documentation states the fine-tuning platform is being wound down and is closed to new users |
| Best fit (OpenAI guidance) | Classification, nuanced translation, format-constrained generation, correcting instruction-following failures | Complex domain-specific tasks requiring advanced reasoning where experts agree on the answer |
Detailed Analysis
Imitation versus optimisation
SFT is imitation. Every gradient step pulls the model toward a response someone has already produced, which makes training stable and predictable and makes the data the product. The method cannot ask for more than the examples contain. It is closely related to knowledge distillation, where the examples come from a stronger model, and to imitation learning generally.
RFT is optimisation against a score. OpenAI's documentation describes the loop: the system "samples several responses per prompt, scores them with the grader, and applies policy-gradient updates based on those rewards". Nothing in that loop requires a reference answer, only a way to tell better from worse. The DeepSeek-R1 paper (January 2025; later published in Nature) argues that this removes the dependence on human-annotated reasoning traces and reports emergent self-reflection, verification and strategy adaptation. Its pure-RL model, R1-Zero, rose from 15.6% to 71.0% on AIME 2024 during training.
The grader is the hard part
RFT moves the difficulty from collecting answers to specifying a reward. A grader that can be satisfied by something other than the intended behaviour will be. OpenAI warns that a model may learn to "reward hack your grader" and that a signal a lucky guess can satisfy is too noisy to train on. DeepSeek chose rule-based accuracy and format rewards for R1-Zero rather than a learned reward model, stating that neural reward models "may suffer from reward hacking" at scale. Tasks with checkable outcomes, such as mathematics, code with tests and structured extraction, suit RFT for this reason. Subjective tasks need rubric graders such as an LLM judge, with the added risk that the judge can be gamed.
Why pipelines use both
Reinforcement learning needs some reward to learn from, so a model that never succeeds gets no signal. SFT is the usual fix. DeepSeek-R1's pipeline has four stages: an SFT cold start on thousands of long chain-of-thought examples, reasoning-focused RL, a second SFT round on about 800,000 samples generated by the RL model, then a final RL stage. The cold start also corrected R1-Zero's readability problems.
A May 2026 write-up by Kim and Ahmad on training a small recursive model reports the same dependency in sharper form: Qwen3.5-4B scored zero on pass@16 before an SFT cold start from a larger teacher, and GRPO then lifted the training reward from 0.3 to 0.6. That is a single task reported in a blog post, and should be weighed accordingly.
Small models and cost
RFT is not automatically the better route to a capable small model. In DeepSeek's comparison at 32B parameters, distilling the large RL-trained model into the small one by SFT (72.6% on AIME 2024) beat applying RL directly (47.0%). RFT also costs more per training prompt, since each needs several sampled completions and grader calls, although published like-for-like cost ratios are scarce. Both methods can be applied to all weights or through adapters such as LoRA; the choice of signal and the choice of which parameters to update are independent.
Best For
Enforcing an output format or house style
Supervised fine-tuningThe target is easy to demonstrate and demonstrations transfer directly.
Classification or extraction with labelled data
Supervised fine-tuningLabels are the supervision; OpenAI lists classification among SFT's core uses.
Reasoning tasks with checkable answers
Reinforcement fine-tuningA programmatic grader gives a clean reward, and RL has been shown to exceed supervised training on verifiable tasks.
Tasks where experts can judge but not easily write the ideal answer
Reinforcement fine-tuningRFT needs a scorer rather than reference outputs, provided expert consensus on quality exists.
Base model that currently fails every attempt
Supervised fine-tuningWith no successes there is no reward signal; an SFT cold start comes first.
Compressing a strong model into a small one
Supervised fine-tuningDistillation by SFT outperformed direct RL at 32B parameters in DeepSeek's comparison.
Agentic tool use in an executable environment
Both / dependsPublished recipes run SFT to stabilise syntax and behaviour, then RL against environment rewards.
No reliable grader available
Supervised fine-tuningA weak or gameable reward invites reward hacking; demonstrations are the safer signal.
The Bottom Line
SFT is the right first tool for most adaptation work. It is simpler, cheaper per example and predictable, and it covers format, style, classification and domain vocabulary. Its limit is that the model learns what the examples show and nothing beyond them.
RFT is worth its extra cost when outcomes can be scored reliably and the goal is to exceed available demonstrations, which is why it sits behind recent reasoning models. It demands a grader that resists exploitation and a model already capable of occasional success. The published pipelines that work best treat the two as stages: supervise to reach competence and readable behaviour, then reinforce to push past the data.
Further Reading
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (January 2025) – arXiv / Nature
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO; February 2024) – arXiv
- Reinforcement fine-tuning guide – OpenAI
- Supervised fine-tuning guide – OpenAI
- Reinforcing Recursive Language Models (Kim and Ahmad, May 2026) – alphaXiv