# Model Merging

> Model merging combines the weights of two or more trained models into one, without further training, so it inherits abilities from each parent.

Source: https://metavert.io/model-merging  
Published: 2026-10-07  
Updated: 2026-10-07

**Model merging** is the practice of combining the weights of two or more trained neural networks into a single model without further training, so that the result inherits abilities from each parent at the inference cost of one. It works almost exclusively between models that share an architecture and a common pretrained ancestor, and it has become a routine tool in the [open-weight](https://metavert.io/open-weight-models) community for composing [fine-tunes](https://metavert.io/fine-tuning).

### Why Averaging Weights Works at All

Averaging two unrelated networks produces noise, because the same function can be encoded by many different arrangements of hidden units. Fine-tunes of one pretrained model are different: they start from the same point and tend to stay in the same low-loss region. *Model soups* (Wortsman et al., March 2022) made the case empirically, showing that averaging the weights of models fine-tuned with different hyperparameters often beat the best single model, with no ensemble cost at inference; a souped ViT-G reached 90.94% top-1 accuracy on ImageNet, a record at the time. For independently trained networks, *Git Re-Basin* (Ainsworth et al., September 2022) showed that hidden units must first be permuted into alignment, and demonstrated this on ResNets trained on CIFAR-10, a much smaller setting than language models.

### The Main Methods

| Method | Idea | Source |
| --- | --- | --- |
| Linear averaging (model soups) | Weighted average of parameters | Wortsman et al., 2022 |
| SLERP | Spherical rather than straight-line interpolation between two models, for a smooth transition between them | Implemented in mergekit; two models only |
| Task arithmetic | Subtract base weights from a fine-tune to get a "task vector"; add, scale or negate vectors on the base | Ilharco et al., December 2022 |
| TIES | Trim small changes, elect a sign per parameter, merge only values that agree with it | Yadav et al., June 2023 |
| DARE | Randomly drop most of each fine-tune's delta and rescale the remainder before merging | Yu et al., November 2023 |
| Passthrough ("frankenmerging") | Stack or splice layers from different models without averaging | mergekit |
| Evolutionary merging | Search automatically over merge recipes in weight space and layer order | Akiba et al., March 2024 |

Task arithmetic reframed merging as editing. Ilharco et al. showed that adding task vectors improved several tasks at once, that negating one reduced performance on its task with little change elsewhere, and that vectors could be combined by analogy to improve a task with no training data for it. TIES addressed why naive addition degrades as more models are combined: redundant small changes and disagreements over a parameter's sign interfere with one another. DARE reported that 90% or even 99% of a fine-tune's delta parameters can be dropped while preserving its abilities, which leaves room for several models' changes to coexist.

### Tooling

Most community merging runs through **mergekit**, an open-source toolkit introduced by Arcee (Goddard et al., March 2024) and licensed under LGPL v3 as of October 2026. Its repository lists more than a dozen methods, among them linear, SLERP, task arithmetic, TIES, DARE and passthrough, and states that merges can run entirely on CPU or with as little as 8 GB of VRAM using an out-of-core approach. It can also extract a [LoRA](https://metavert.io/lora)-style low-rank approximation from a fine-tuned model, and adapters can themselves be combined: the [Hugging Face](https://metavert.io/hugging-face) PEFT library supports weighted combinations of multiple LoRA adapters. The cost profile is the appeal. A merge takes minutes on commodity hardware, against the GPU-hours of a training run.

### What Works and What Breaks

**Works:** averaging several fine-tunes of the same base on the same task for robustness; combining a small number of fine-tunes of one base with complementary skills; blending adapters. The evolutionary merging paper reported a Japanese-language model with mathematical reasoning ability built from separate Japanese and math models, a result its authors described as surprising.

**Breaks:** merging models with different architectures or unrelated training lineages; merging many models at once, where the interference TIES describes accumulates; merging fine-tunes that have drifted far from their base. Results are also hard to predict. Recipes are found by trial, evaluation and intuition, which is the gap evolutionary search tries to fill, and a merge tuned against a public leaderboard can overfit to it like any other heavily searched configuration. The mergekit paper cites leaderboard standing as evidence of merged models' strength, which is a reason to test a merge on private, task-specific [evaluations](https://metavert.io/llm-evaluation) before trusting it.

### Licensing and Provenance

A merged model is a derivative of every parent, and the parents may carry different licenses, some with non-commercial or use-based restrictions. Hugging Face model cards allow a merge to declare a list of base models in metadata, which the Hub displays as a merge relationship, and a license field; both are self-reported by the uploader and not verified. In practice, establishing whether a merged checkpoint can be used commercially means tracing each ancestor's terms, and merges of merges make that chain long. Provenance matters for safety as well: a merge inherits whatever its parents learned, including behavior introduced by [training data](https://metavert.io/ai-training-data) no one involved in the merge has inspected. For regulated or commercial deployments, an undocumented lineage is a reason to prefer a model with a single known origin.

## Related Topics

- [Fine-Tuning](https://metavert.io/fine-tuning) — Produces the specialized models that get merged
- [LoRA (Low-Rank Adaptation)](https://metavert.io/lora) — Adapters can be merged, blended and extracted
- [Open-Weight Models](https://metavert.io/open-weight-models) — Merging requires access to weights
- [Hugging Face](https://metavert.io/hugging-face) — Where merges are published and their lineage declared
- [Knowledge Distillation](https://metavert.io/knowledge-distillation) — The training-based way to combine models
- [Continual Learning](https://metavert.io/continual-learning) — Merging as a way to add skills without forgetting
- [Parameter-Efficient Fine-Tuning](https://metavert.io/parameter-efficient-fine-tuning) — Small deltas that compose more easily
- [LLM Evaluation](https://metavert.io/llm-evaluation) — The only way to know whether a merge worked
- [Small Language Models](https://metavert.io/small-language-models) — Common targets for community merging

## Further Reading

- [Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time](https://arxiv.org/abs/2203.05482) — Wortsman et al., arXiv, March 2022
- [Git Re-Basin: Merging Models modulo Permutation Symmetries](https://arxiv.org/abs/2209.04836) — Ainsworth et al., arXiv, September 2022
- [Editing Models with Task Arithmetic](https://arxiv.org/abs/2212.04089) — Ilharco et al., arXiv, December 2022
- [TIES-Merging: Resolving Interference When Merging Models](https://arxiv.org/abs/2306.01708) — Yadav et al., arXiv, June 2023
- [Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch](https://arxiv.org/abs/2311.03099) — Yu et al., arXiv, November 2023
- [Arcee's MergeKit: A Toolkit for Merging Large Language Models](https://arxiv.org/abs/2403.13257) — Goddard et al., arXiv, March 2024
- [mergekit repository](https://github.com/arcee-ai/mergekit) — Arcee AI on GitHub, accessed October 2026
- [Evolutionary Optimization of Model Merging Recipes](https://arxiv.org/abs/2403.13187) — Akiba et al., arXiv, March 2024
- [Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities](https://arxiv.org/abs/2408.07666) — Yang et al., arXiv, August 2024
- [Model Cards (base model and license metadata)](https://huggingface.co/docs/hub/model-cards) — Hugging Face Hub documentation, accessed October 2026
