# Continual Learning

> Continual learning is a model's ability to keep learning from new data over time without catastrophically forgetting what it already knows.

Source: https://metavert.io/continual-learning  
Published: 2026-10-07  
Updated: 2026-10-07

**Continual learning** is the ability of a machine learning system to keep acquiring knowledge and skills from new data over time without losing what it already knows. The central obstacle is *catastrophic forgetting*: training a neural network on a new task tends to overwrite the weights that encoded earlier ones. For deployed [large language models](https://metavert.io/large-language-models), whose weights are frozen at release, the term also names an unsolved gap: the model that answers on day 500 has learned nothing since day one.

### The Stability-Plasticity Problem

A network must be plastic enough to absorb new information and stable enough to retain the old, and gradient descent on new data optimizes only the first. A widely cited survey (Wang et al., January 2023, later in IEEE TPAMI) summarizes the field's objective as a proper stability-plasticity trade-off together with good generalization within and across tasks, under resource constraints. That last clause matters: retraining from scratch on all data ever seen solves forgetting trivially and is exactly what continual learning tries to avoid.

### Classic Approaches

| Family | Idea | Example | Cost |
| --- | --- | --- | --- |
| Replay | Keep or regenerate samples of old data and mix them into new training | Gradient Episodic Memory (Lopez-Paz and Ranzato, 2017) stores examples and constrains updates so past-task loss does not rise | Storage, and access to old data that may no longer be retainable |
| Regularization | Penalize changes to weights that mattered for earlier tasks | Elastic Weight Consolidation (Kirkpatrick et al., 2016) selectively slows learning on important weights | Importance estimates degrade as tasks accumulate |
| Parameter isolation | Give each task its own parameters and freeze the rest | Progressive Neural Networks (Rusu et al., 2016) add a new column per task with lateral connections | Model grows with every task |

These results were demonstrated at small scale: EWC on MNIST variants and sequences of Atari games, GEM on MNIST and CIFAR-100 variants, progressive networks on Atari and 3D mazes. A 2019 comparison of eleven methods (De Lange et al.) found outcomes sensitive to model capacity, regularization and even the order in which tasks arrive. None of these techniques became a standard component of large-model training.

### What LLM-Era Systems Do Instead

Production systems mostly sidestep the problem by keeping the weights fixed and moving what must change outside them.

**Retrieval.** The original [retrieval-augmented generation](https://metavert.io/retrieval-augmented-generation) paper (Lewis et al., May 2020) was motivated in part by the observation that updating a model's world knowledge remained an open research problem; pairing parametric memory with a searchable index lets knowledge be updated by editing documents.

**Memory files.** Agents write notes to persistent storage and read them back later. Anthropic's memory tool, for example, lets a model create, read, update and delete files in a memory directory that persists between sessions, with storage controlled by the developer's application. Project instruction files such as [CLAUDE.md](https://metavert.io/CLAUDE.md) serve the same purpose with a human editor. This is [agentic memory](https://metavert.io/agentic-memory): learning as record-keeping, bounded by what fits back into the [context window](https://metavert.io/context-windows).

**Periodic fine-tunes and new releases.** Accumulated data is folded into the weights in batches, through [fine-tuning](https://metavert.io/fine-tuning) or the next model version. A survey of continual learning for LLMs (Shi et al., April 2024) organizes this into continual pretraining, domain-adaptive pretraining and continual fine-tuning, and notes that models tailored to specific needs often degrade on previously known domains.

**Adapters.** [Parameter-efficient methods](https://metavert.io/parameter-efficient-fine-tuning) are parameter isolation in modern form. Houlsby et al. (2019) noted that with adapters new tasks can be added without revisiting previous ones, and Biderman et al. (May 2024) found that [LoRA](https://metavert.io/lora) preserved more out-of-domain performance than full fine-tuning in two domains. [Model merging](https://metavert.io/model-merging) is also studied as a way to combine separately trained skills.

### Where the Research Stands

As of October 2026, no widely deployed language model updates its own weights from individual interactions in production; published deployments rely on the workarounds above. Each has a ceiling. Retrieval and memory files supply facts and instructions but do not obviously build skill: the model reading its notes has the same underlying competence as before. Periodic fine-tuning learns in coarse steps and risks forgetting each time. Adapters isolate tasks but do not integrate them.

Whether in-weights continual learning is necessary for more capable agents, or whether larger context, better memory tooling and frequent [post-training](https://metavert.io/post-training) are sufficient, is a live disagreement rather than a settled question. The classic literature offers mechanisms, but its evidence comes from benchmarks far smaller than current models, and the LLM-specific literature is still largely organized around surveys and taxonomies rather than a method that practitioners have adopted. Claims that the problem is solved, or nearly so, should be checked against what is actually running in production.

## Related Topics

- [Agentic Memory](https://metavert.io/agentic-memory) — How agents remember without changing weights
- [Retrieval-Augmented Generation](https://metavert.io/retrieval-augmented-generation) — Updating knowledge by updating documents
- [Fine-Tuning](https://metavert.io/fine-tuning) — The batch route for getting new data into weights
- [Parameter-Efficient Fine-Tuning](https://metavert.io/parameter-efficient-fine-tuning) — Parameter isolation in modern form
- [LoRA (Low-Rank Adaptation)](https://metavert.io/lora) — Adapters that forget less than full fine-tuning
- [Model Merging](https://metavert.io/model-merging) — Combining separately learned skills in weight space
- [Post-Training](https://metavert.io/post-training) — Where periodic updates actually happen
- [Context Windows](https://metavert.io/context-windows) — The limit on memory-as-text
- [Transfer Learning](https://metavert.io/transfer-learning) — Reusing knowledge across tasks

## Further Reading

- [Overcoming catastrophic forgetting in neural networks](https://arxiv.org/abs/1612.00796) — Kirkpatrick et al., arXiv, December 2016
- [Gradient Episodic Memory for Continual Learning](https://arxiv.org/abs/1706.08840) — Lopez-Paz and Ranzato, arXiv, June 2017
- [Progressive Neural Networks](https://arxiv.org/abs/1606.04671) — Rusu et al., arXiv, June 2016
- [A continual learning survey: Defying forgetting in classification tasks](https://arxiv.org/abs/1909.08383) — De Lange et al., arXiv, September 2019
- [A Comprehensive Survey of Continual Learning: Theory, Method and Application](https://arxiv.org/abs/2302.00487) — Wang et al., arXiv, January 2023
- [Continual Learning of Large Language Models: A Comprehensive Survey](https://arxiv.org/abs/2404.16789) — Shi et al., arXiv, April 2024
- [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://arxiv.org/abs/2005.11401) — Lewis et al., arXiv, May 2020
- [Memory tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool) — Anthropic documentation, accessed October 2026
- [Parameter-Efficient Transfer Learning for NLP](https://arxiv.org/abs/1902.00751) — Houlsby et al., arXiv, February 2019
- [LoRA Learns Less and Forgets Less](https://arxiv.org/abs/2405.09673) — Biderman et al., arXiv, May 2024
