LoRA vs Full Fine-Tuning

Comparison

LoRA (Low-Rank Adaptation) and full fine-tuning are two ways to adapt a pretrained model to a task. Full fine-tuning updates every weight. LoRA, the most widely used form of parameter-efficient fine-tuning, freezes the pretrained weights and trains a pair of small matrices whose product is added to selected weight matrices.

LoRA is far cheaper to train, store and swap, and on many tasks it matches full fine-tuning. It is not a free equivalent. Later studies found that at commonly used ranks LoRA learns less on hard domain shifts, forgets less of what the base model knew, and reaches solutions that are structurally different from those of full fine-tuning. Which matters more depends on how far the target task is from the base model's training.

Feature Comparison

DimensionLoRAFull fine-tuning
What is trainedLow-rank matrices B and A added to frozen weights (W0 + BA)All model parameters
Trainable parametersUp to 10,000× fewer on GPT-3 175B (Hu et al., 2021)100% of parameters
Training memoryGPT-3 175B: 350 GB of VRAM (Hu et al.)GPT-3 175B: 1.2 TB of VRAM (same source)
With a quantized base (QLoRA)65B-parameter model on a single 48 GB GPU (Dettmers et al., 2023)More than 780 GB of GPU memory for 16-bit fine-tuning of the same model
Checkpoint size per taskAbout 35 MB for GPT-3 at rank 4About 350 GB for GPT-3
Inference latencyNone added once the matrices are merged into the base weightsBaseline
Serving many tasksOne shared base model with swappable adapters; merged weights cannot batch different tasks in one passA separate full copy of the model per task
Quality on 2021 NLU/NLG benchmarksOn par or better on RoBERTa, DeBERTa, GPT-2 and GPT-3 (Hu et al.)Baseline
Quality on code, continued pretraining (Llama-2-7B)HumanEval 0.224 at rank 256 (Biderman et al., 2024)HumanEval 0.263
Quality on code, instruction tuning (Llama-2-7B)HumanEval 0.358 at rank 16; 0.498 at rank 256HumanEval 0.497
Forgetting of base-model abilitiesLess forgetting than full fine-tuning, and less than weight decay or dropout (Biderman et al.)More forgetting outside the target domain
Rank of the learned updateFixed by configuration (commonly 4 to 256)10–100× higher than typical LoRA ranks (Biderman et al.)

Detailed Analysis

Why LoRA is cheap

Hu et al. (Microsoft, June 2021) proposed that the change in weights during adaptation has low intrinsic rank, and so can be represented as the product of two thin matrices. Only those matrices receive gradients. Because the optimizer no longer stores state for the frozen weights, the paper reports a reduction in training VRAM on GPT-3 175B from 1.2 TB to 350 GB, a 10,000-fold reduction in checkpoint size at rank 4 with only the query and value projections adapted, and a 25% training speed-up on that model. After training, the product can be added back into the base weights, so a merged LoRA model runs with no additional inference latency.

QLoRA (Dettmers et al., May 2023) extended the saving by back-propagating through a frozen base model quantized to 4 bits (see model quantization). The authors report fine-tuning a 65B model on one 48 GB GPU, where 16-bit full fine-tuning would need more than 780 GB. Their direct comparison with full fine-tuning covers models up to 3B parameters; at 33B and 65B the full fine-tuning baseline was too expensive to run, so equivalence at that scale is argued rather than measured.

Does it match full fine-tuning?

The original paper says yes, on GLUE and on generation tasks such as WikiSQL and SAMSum. Those are modest shifts for the models involved. Biderman et al. ("LoRA Learns Less and Forgets Less", TMLR 2024) tested harder shifts on Llama-2-7B, in programming and mathematics, under both instruction tuning (about 100K examples) and continued pretraining (20B tokens). Their conclusion: "in the standard low-rank settings, LoRA substantially underperforms full finetuning".

The detail matters. For code instruction tuning, rank 16 reached 0.358 on HumanEval against 0.497 for full fine-tuning, while rank 256 reached 0.498 and closed the gap. For continued pretraining on code the gap persisted at every rank tested (0.224 vs 0.263). For mathematics instruction tuning, a smaller shift from the base model, the gap was small (0.634 vs 0.642 on GSM8K at rank 256). These results come from one model family and two domains. QLoRA's authors made a related observation: adapting only the query and value matrices was not enough to match full fine-tuning on larger models, and adapters on all linear layers were needed.

Forgetting and structural differences

The same constraint that limits learning also protects the base model. Biderman et al. found LoRA retained more performance outside the target domain than full fine-tuning and more than standard regularizers, and kept generations more diverse. For teams that want a specialist that still behaves like the base model elsewhere, that is a benefit and not just a consolation.

Shuttleworth et al. ("An Illusion of Equivalence", 2024–2025) add a caution. Even where task accuracy matches, LoRA-trained weights contain new high-ranking singular vectors, which the authors call intruder dimensions, that are absent after full fine-tuning. They link these to forgetting and report that models fine-tuned sequentially with LoRA accumulate them and tend to perform worse in continual learning settings. Equal benchmark scores do not imply equal models.

Tuning and speed in practice

LoRA has its own sensitivities. Biderman et al. report that it needs learning rates about an order of magnitude higher than full fine-tuning (roughly 5e-5 to 5e-4), that setting alpha to twice the rank matters at high ranks, and that targeting all modules, MLP blocks included, beats attention-only adapters. They also note that in standard implementations LoRA tended to train more slowly than full fine-tuning at a fixed batch size, which contrasts with the 25% speed-up reported for GPT-3 in the original paper. The memory saving is consistent across sources; the wall-clock saving is not.

Best For

Fine-tuning on a single consumer or workstation GPU

LoRA

With a 4-bit base, QLoRA fits a 65B model in 48 GB; full fine-tuning of the same model needs more than 780 GB.

Many task- or customer-specific variants of one model

LoRA

Adapters are megabytes instead of full copies, and can be swapped over a shared base.

Large domain shift through continued pretraining

Full fine-tuning

LoRA trailed full fine-tuning at every rank tested for code continued pretraining in Biderman et al.

Instruction tuning for a new format or style

LoRA

Small shifts are where LoRA matches full fine-tuning, and higher ranks close most remaining gaps.

Keeping the base model's general abilities intact

LoRA

LoRA forgot less outside the target domain than full fine-tuning or standard regularizers.

Maximum task quality with compute available

Full fine-tuning

Full fine-tuning learns updates of 10–100× higher rank and remains the upper reference in the comparative studies.

Sequential updates to the same model over time

Both / depends

LoRA forgets less per step, but repeated LoRA fine-tuning accumulated intruder dimensions and performed worse in the continual-learning tests of Shuttleworth et al.

Latency-critical serving of one fixed task

Both / depends

A merged LoRA adapter adds no latency, so serving cost is the same either way.

The Bottom Line

LoRA is the sensible default for most adaptation: instruction tuning, style and format, and any setting where hardware, storage or the number of variants is the binding constraint. The memory savings are large and well documented, and merged adapters cost nothing at inference.

Full fine-tuning remains the reference when the target domain is far from what the base model knows, when training on billions of tokens of new material, or when the last points of task quality justify the compute. The evidence against treating the two as interchangeable is solid but narrow: the main comparative studies use 7B-class models in two domains. A fair summary is the title of the best-known one. LoRA learns less and forgets less, and raising the rank and adapting all layers trades the second property for the first.