Supervised Fine-Tuning vs Reinforcement Fine-Tuning

Comparison

Supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) are the two main ways to change a pretrained model's behaviour during post-training. SFT, the default meaning of fine-tuning, shows the model example outputs and trains it to reproduce them. RFT lets the model produce its own outputs, scores them with a grader, and shifts the weights toward higher-scoring behaviour, typically with a policy-gradient method such as GRPO.

The practical rule is short. Use SFT when good answers can be written down. Use RFT when good answers can be recognised and scored but not easily demonstrated, and when the model can already succeed some of the time. The strongest published systems use both, in sequence.

Feature Comparison

DimensionSupervised fine-tuningReinforcement fine-tuning
Training signalExample outputs: the model is trained to maximise the likelihood of demonstrated responsesA numeric reward from a grader applied to the model's own sampled outputs
What must be suppliedPrompt–response pairsPrompts plus a grader (programmatic check, rubric or model-based scorer)
Typical algorithmCross-entropy on target tokensPolicy gradient; GRPO, a PPO variant introduced in DeepSeekMath (February 2024), is widely used
Data volume (OpenAI guidance)Minimum 10 examples; improvements typically seen from 50–100Start with several dozen to a few hundred; up to 50,000 training examples
Performance ceilingTied to the quality of the demonstrationsCan exceed the demonstrations: DeepSeek-R1 reports surpassing counterparts trained by supervised learning on human demonstrations for verifiable tasks
Starting requirementNone beyond a base modelThe model must sometimes succeed: OpenAI advises against RFT when scores sit at the minimum or maximum
Characteristic failureReproduces the limits and errors of the examplesReward hacking: scoring well on the grader without being correct (noted in OpenAI's documentation and by DeepSeek)
Output styleFollows the demonstrationsCan drift: DeepSeek-R1-Zero showed poor readability and language mixing
Small-model evidenceDistillation by SFT: DeepSeek-R1-Distill-Qwen-32B scored 72.6% on AIME 2024RL applied directly at the same size: DeepSeek-R1-Zero-Qwen-32B scored 47.0%
Compute per exampleOne forward and backward pass per exampleSeveral sampled responses per prompt, each scored; exact cost ratios are not publicly documented
Hosted availability (OpenAI, October 2026)GPT-4.1, GPT-4.1 mini and GPT-4.1 nanoo4-mini only; OpenAI's documentation states the fine-tuning platform is being wound down and is closed to new users
Best fit (OpenAI guidance)Classification, nuanced translation, format-constrained generation, correcting instruction-following failuresComplex domain-specific tasks requiring advanced reasoning where experts agree on the answer

Detailed Analysis

Imitation versus optimisation

SFT is imitation. Every gradient step pulls the model toward a response someone has already produced, which makes training stable and predictable and makes the data the product. The method cannot ask for more than the examples contain. It is closely related to knowledge distillation, where the examples come from a stronger model, and to imitation learning generally.

RFT is optimisation against a score. OpenAI's documentation describes the loop: the system "samples several responses per prompt, scores them with the grader, and applies policy-gradient updates based on those rewards". Nothing in that loop requires a reference answer, only a way to tell better from worse. The DeepSeek-R1 paper (January 2025; later published in Nature) argues that this removes the dependence on human-annotated reasoning traces and reports emergent self-reflection, verification and strategy adaptation. Its pure-RL model, R1-Zero, rose from 15.6% to 71.0% on AIME 2024 during training.

The grader is the hard part

RFT moves the difficulty from collecting answers to specifying a reward. A grader that can be satisfied by something other than the intended behaviour will be. OpenAI warns that a model may learn to "reward hack your grader" and that a signal a lucky guess can satisfy is too noisy to train on. DeepSeek chose rule-based accuracy and format rewards for R1-Zero rather than a learned reward model, stating that neural reward models "may suffer from reward hacking" at scale. Tasks with checkable outcomes, such as mathematics, code with tests and structured extraction, suit RFT for this reason. Subjective tasks need rubric graders such as an LLM judge, with the added risk that the judge can be gamed.

Why pipelines use both

Reinforcement learning needs some reward to learn from, so a model that never succeeds gets no signal. SFT is the usual fix. DeepSeek-R1's pipeline has four stages: an SFT cold start on thousands of long chain-of-thought examples, reasoning-focused RL, a second SFT round on about 800,000 samples generated by the RL model, then a final RL stage. The cold start also corrected R1-Zero's readability problems.

A May 2026 write-up by Kim and Ahmad on training a small recursive model reports the same dependency in sharper form: Qwen3.5-4B scored zero on pass@16 before an SFT cold start from a larger teacher, and GRPO then lifted the training reward from 0.3 to 0.6. That is a single task reported in a blog post, and should be weighed accordingly.

Small models and cost

RFT is not automatically the better route to a capable small model. In DeepSeek's comparison at 32B parameters, distilling the large RL-trained model into the small one by SFT (72.6% on AIME 2024) beat applying RL directly (47.0%). RFT also costs more per training prompt, since each needs several sampled completions and grader calls, although published like-for-like cost ratios are scarce. Both methods can be applied to all weights or through adapters such as LoRA; the choice of signal and the choice of which parameters to update are independent.

Best For

Enforcing an output format or house style

Supervised fine-tuning

The target is easy to demonstrate and demonstrations transfer directly.

Classification or extraction with labelled data

Supervised fine-tuning

Labels are the supervision; OpenAI lists classification among SFT's core uses.

Reasoning tasks with checkable answers

Reinforcement fine-tuning

A programmatic grader gives a clean reward, and RL has been shown to exceed supervised training on verifiable tasks.

Tasks where experts can judge but not easily write the ideal answer

Reinforcement fine-tuning

RFT needs a scorer rather than reference outputs, provided expert consensus on quality exists.

Base model that currently fails every attempt

Supervised fine-tuning

With no successes there is no reward signal; an SFT cold start comes first.

Compressing a strong model into a small one

Supervised fine-tuning

Distillation by SFT outperformed direct RL at 32B parameters in DeepSeek's comparison.

Agentic tool use in an executable environment

Both / depends

Published recipes run SFT to stabilise syntax and behaviour, then RL against environment rewards.

No reliable grader available

Supervised fine-tuning

A weak or gameable reward invites reward hacking; demonstrations are the safer signal.

The Bottom Line

SFT is the right first tool for most adaptation work. It is simpler, cheaper per example and predictable, and it covers format, style, classification and domain vocabulary. Its limit is that the model learns what the examples show and nothing beyond them.

RFT is worth its extra cost when outcomes can be scored reliably and the goal is to exceed available demonstrations, which is why it sits behind recent reasoning models. It demands a grader that resists exploitation and a model already capable of occasional success. The published pipelines that work best treat the two as stages: supervise to reach competence and readable behaviour, then reinforce to push past the data.