# Context Windows

> Context windows define how much information an AI model can process at once, with frontier models now handling 100K-2M tokens in a single pass.

Source: https://metavert.io/context-windows  
Updated: 2026-04-29

A **context window** is the maximum amount of text (measured in tokens) that a [large language model](https://metavert.io/large-language-models) can process in a single interaction—encompassing both the input (prompt, documents, conversation history) and the output (the model's response). It is one of the most practically significant parameters defining what an AI system can do.

Context windows have expanded dramatically. GPT-3 (2020) supported 4,096 tokens—roughly 3,000 words. By 2024, Claude offered 200,000 tokens and Gemini reached 1 million. In 2025-2026, experimental models push toward 2 million tokens and beyond. This expansion transforms capability: a model with a 4K context can answer a question about a paragraph; a model with a 200K context can analyze an entire codebase, legal contract, or research paper in one pass.

The engineering behind long contexts is non-trivial. The [self-attention mechanism](https://metavert.io/attention-mechanism) scales quadratically with context length—doubling the context quadruples compute and memory requirements. Innovations like Flash Attention (IO-aware attention computation), RoPE (Rotary Position Embeddings for position encoding), sliding window attention, and KV-cache optimization have made long contexts practical. [High Bandwidth Memory (HBM)](https://metavert.io/high-bandwidth-memory) is critical because the key-value cache grows linearly with context length and must remain in fast memory.

For [AI agents](https://metavert.io/agentic-ai), context windows define working memory. An agent operating on 14-hour autonomous task horizons must maintain context about what it's doing, what it's already tried, and what it's learned. Combined with [RAG](https://metavert.io/retrieval-augmented-generation) (which extends effective context by retrieving relevant information from external stores) and [vector search](https://metavert.io/vector-search), long context windows enable agents to work on complex projects that require understanding thousands of interconnected details simultaneously.

## Related Topics

- [Context Engineering](https://metavert.io/context-engineering) — The discipline of managing what fills the window
- [Large Language Models](https://metavert.io/large-language-models)
- [Attention Mechanism](https://metavert.io/attention-mechanism)
- [Transformer Architecture](https://metavert.io/transformer-architecture)
- [High Bandwidth Memory](https://metavert.io/high-bandwidth-memory)
- [Retrieval Augmented Generation](https://metavert.io/retrieval-augmented-generation)
- [Agentic Memory](https://metavert.io/agentic-memory)
- [AI Inference](https://metavert.io/ai-inference)

## Further Reading

- [The State of AI Agents in 2026](https://meditations.metavert.io/p/the-state-of-ai-agents-in-2026)
