Model Merging

Model merging is the practice of combining the weights of two or more trained neural networks into a single model without further training, so that the result inherits abilities from each parent at the inference cost of one. It works almost exclusively between models that share an architecture and a common pretrained ancestor, and it has become a routine tool in the open-weight community for composing fine-tunes.

Why Averaging Weights Works at All

Averaging two unrelated networks produces noise, because the same function can be encoded by many different arrangements of hidden units. Fine-tunes of one pretrained model are different: they start from the same point and tend to stay in the same low-loss region. Model soups (Wortsman et al., March 2022) made the case empirically, showing that averaging the weights of models fine-tuned with different hyperparameters often beat the best single model, with no ensemble cost at inference; a souped ViT-G reached 90.94% top-1 accuracy on ImageNet, a record at the time. For independently trained networks, Git Re-Basin (Ainsworth et al., September 2022) showed that hidden units must first be permuted into alignment, and demonstrated this on ResNets trained on CIFAR-10, a much smaller setting than language models.

The Main Methods

MethodIdeaSource
Linear averaging (model soups)Weighted average of parametersWortsman et al., 2022
SLERPSpherical rather than straight-line interpolation between two models, for a smooth transition between themImplemented in mergekit; two models only
Task arithmeticSubtract base weights from a fine-tune to get a "task vector"; add, scale or negate vectors on the baseIlharco et al., December 2022
TIESTrim small changes, elect a sign per parameter, merge only values that agree with itYadav et al., June 2023
DARERandomly drop most of each fine-tune's delta and rescale the remainder before mergingYu et al., November 2023
Passthrough ("frankenmerging")Stack or splice layers from different models without averagingmergekit
Evolutionary mergingSearch automatically over merge recipes in weight space and layer orderAkiba et al., March 2024

Task arithmetic reframed merging as editing. Ilharco et al. showed that adding task vectors improved several tasks at once, that negating one reduced performance on its task with little change elsewhere, and that vectors could be combined by analogy to improve a task with no training data for it. TIES addressed why naive addition degrades as more models are combined: redundant small changes and disagreements over a parameter's sign interfere with one another. DARE reported that 90% or even 99% of a fine-tune's delta parameters can be dropped while preserving its abilities, which leaves room for several models' changes to coexist.

Tooling

Most community merging runs through mergekit, an open-source toolkit introduced by Arcee (Goddard et al., March 2024) and licensed under LGPL v3 as of October 2026. Its repository lists more than a dozen methods, among them linear, SLERP, task arithmetic, TIES, DARE and passthrough, and states that merges can run entirely on CPU or with as little as 8 GB of VRAM using an out-of-core approach. It can also extract a LoRA-style low-rank approximation from a fine-tuned model, and adapters can themselves be combined: the Hugging Face PEFT library supports weighted combinations of multiple LoRA adapters. The cost profile is the appeal. A merge takes minutes on commodity hardware, against the GPU-hours of a training run.

What Works and What Breaks

Works: averaging several fine-tunes of the same base on the same task for robustness; combining a small number of fine-tunes of one base with complementary skills; blending adapters. The evolutionary merging paper reported a Japanese-language model with mathematical reasoning ability built from separate Japanese and math models, a result its authors described as surprising.

Breaks: merging models with different architectures or unrelated training lineages; merging many models at once, where the interference TIES describes accumulates; merging fine-tunes that have drifted far from their base. Results are also hard to predict. Recipes are found by trial, evaluation and intuition, which is the gap evolutionary search tries to fill, and a merge tuned against a public leaderboard can overfit to it like any other heavily searched configuration. The mergekit paper cites leaderboard standing as evidence of merged models' strength, which is a reason to test a merge on private, task-specific evaluations before trusting it.

Licensing and Provenance

A merged model is a derivative of every parent, and the parents may carry different licenses, some with non-commercial or use-based restrictions. Hugging Face model cards allow a merge to declare a list of base models in metadata, which the Hub displays as a merge relationship, and a license field; both are self-reported by the uploader and not verified. In practice, establishing whether a merged checkpoint can be used commercially means tracing each ancestor's terms, and merges of merges make that chain long. Provenance matters for safety as well: a merge inherits whatever its parents learned, including behavior introduced by training data no one involved in the merge has inspected. For regulated or commercial deployments, an undocumented lineage is a reason to prefer a model with a single known origin.

Further Reading