Continual Learning

Continual learning is the ability of a machine learning system to keep acquiring knowledge and skills from new data over time without losing what it already knows. The central obstacle is catastrophic forgetting: training a neural network on a new task tends to overwrite the weights that encoded earlier ones. For deployed large language models, whose weights are frozen at release, the term also names an unsolved gap: the model that answers on day 500 has learned nothing since day one.

The Stability-Plasticity Problem

A network must be plastic enough to absorb new information and stable enough to retain the old, and gradient descent on new data optimizes only the first. A widely cited survey (Wang et al., January 2023, later in IEEE TPAMI) summarizes the field's objective as a proper stability-plasticity trade-off together with good generalization within and across tasks, under resource constraints. That last clause matters: retraining from scratch on all data ever seen solves forgetting trivially and is exactly what continual learning tries to avoid.

Classic Approaches

FamilyIdeaExampleCost
ReplayKeep or regenerate samples of old data and mix them into new trainingGradient Episodic Memory (Lopez-Paz and Ranzato, 2017) stores examples and constrains updates so past-task loss does not riseStorage, and access to old data that may no longer be retainable
RegularizationPenalize changes to weights that mattered for earlier tasksElastic Weight Consolidation (Kirkpatrick et al., 2016) selectively slows learning on important weightsImportance estimates degrade as tasks accumulate
Parameter isolationGive each task its own parameters and freeze the restProgressive Neural Networks (Rusu et al., 2016) add a new column per task with lateral connectionsModel grows with every task

These results were demonstrated at small scale: EWC on MNIST variants and sequences of Atari games, GEM on MNIST and CIFAR-100 variants, progressive networks on Atari and 3D mazes. A 2019 comparison of eleven methods (De Lange et al.) found outcomes sensitive to model capacity, regularization and even the order in which tasks arrive. None of these techniques became a standard component of large-model training.

What LLM-Era Systems Do Instead

Production systems mostly sidestep the problem by keeping the weights fixed and moving what must change outside them.

Retrieval. The original retrieval-augmented generation paper (Lewis et al., May 2020) was motivated in part by the observation that updating a model's world knowledge remained an open research problem; pairing parametric memory with a searchable index lets knowledge be updated by editing documents.

Memory files. Agents write notes to persistent storage and read them back later. Anthropic's memory tool, for example, lets a model create, read, update and delete files in a memory directory that persists between sessions, with storage controlled by the developer's application. Project instruction files such as CLAUDE.md serve the same purpose with a human editor. This is agentic memory: learning as record-keeping, bounded by what fits back into the context window.

Periodic fine-tunes and new releases. Accumulated data is folded into the weights in batches, through fine-tuning or the next model version. A survey of continual learning for LLMs (Shi et al., April 2024) organizes this into continual pretraining, domain-adaptive pretraining and continual fine-tuning, and notes that models tailored to specific needs often degrade on previously known domains.

Adapters. Parameter-efficient methods are parameter isolation in modern form. Houlsby et al. (2019) noted that with adapters new tasks can be added without revisiting previous ones, and Biderman et al. (May 2024) found that LoRA preserved more out-of-domain performance than full fine-tuning in two domains. Model merging is also studied as a way to combine separately trained skills.

Where the Research Stands

As of October 2026, no widely deployed language model updates its own weights from individual interactions in production; published deployments rely on the workarounds above. Each has a ceiling. Retrieval and memory files supply facts and instructions but do not obviously build skill: the model reading its notes has the same underlying competence as before. Periodic fine-tuning learns in coarse steps and risks forgetting each time. Adapters isolate tasks but do not integrate them.

Whether in-weights continual learning is necessary for more capable agents, or whether larger context, better memory tooling and frequent post-training are sufficient, is a live disagreement rather than a settled question. The classic literature offers mechanisms, but its evidence comes from benchmarks far smaller than current models, and the LLM-specific literature is still largely organized around surveys and taxonomies rather than a method that practitioners have adopted. Claims that the problem is solved, or nearly so, should be checked against what is actually running in production.

Further Reading