# AI Model Pricing

> AI model pricing is how hosted LLMs are billed per token: input, output, cached and reasoning tokens, batch discounts, and list prices as of October 2026.

Source: https://metavert.io/ai-model-pricing  
Published: 2026-10-07  
Updated: 2026-10-07

**AI model pricing** is the way providers of hosted language models charge for use: almost always per token, with separate rates for the text a model reads (input), the text it writes (output), and input it has seen before (cached), plus discounts or surcharges for how and when the request is run. Understanding the structure matters more than memorising any price, because the list price per token is only loosely related to what a task costs.

### How API Pricing Works

Prices are quoted per million tokens. Five mechanisms account for most of a bill.

**Input and output.** Output tokens cost more than input tokens — five times more on every Anthropic and OpenAI flagship row checked for this page. Agent workloads are input-heavy, because the whole conversation, tool definitions and tool results are resent on every turn.

**Cached input.** Providers discount input that repeats a previously processed prefix. Anthropic charges a premium to write a cache entry (1.25x the base input price for a five-minute entry, 2x for one hour) and 0.1x to read it, falling to 0.05x on Claude Opus 5.5 and 0.025x on Claude Fable 5.1. OpenAI lists a separate cached-input rate per model. DeepSeek enables prefix caching by default. Because agents resend a stable prefix, cache behaviour often dominates real cost, and anything that rewrites the prefix — including [context compaction](https://metavert.io/context-compaction) — invalidates it.

**Batch.** Anthropic, OpenAI and Google all offer roughly a 50% discount for asynchronous batch processing, which suits evaluation runs and offline data generation but not interactive agents.

**Reasoning tokens.** [Reasoning models](https://metavert.io/reasoning-models) generate hidden thinking before answering. OpenAI's documentation states that reasoning tokens "are billed as output tokens" and occupy context even though they are not returned; Google's price tables quote output "including thinking tokens". A higher reasoning-effort setting therefore raises cost without changing the visible answer length (see [test-time compute](https://metavert.io/test-time-compute)).

**Modifiers.** Long prompts can cost more: OpenAI applies higher rates above 272K input tokens and Google above 200K on Gemini 3.1 Pro Preview, while Anthropic bills its 1M-token window at the standard rate. Faster service tiers carry premiums, and DeepSeek halves its prices off-peak.

### List Prices as of October 2026

The figures below were read from each vendor's own pricing page on 6 October 2026. They are standard-tier, short-context list prices in US dollars per million tokens for a selection of models.

| Vendor | Model | Input | Cached input | Output |
| --- | --- | --- | --- | --- |
| Anthropic | Claude Fable 5.1 | $10.00 | $0.25 | $50.00 |
| Anthropic | Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
| Anthropic | Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 |
| Anthropic | Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| OpenAI | gpt-6-astra | $10.00 | $1.00 | $50.00 |
| OpenAI | gpt-6.1-sol | $2.00 | $0.10 | $10.00 |
| OpenAI | gpt-6-luna | $0.10 | $0.01 | $0.50 |
| Google | Gemini 3.1 Pro Preview (prompts up to 200K) | $2.00 | $0.20 | $12.00 |
| Google | Gemini 3.8 Flash (through 31 December 2026) | $0.75 | $0.075 | $3.75 |
| DeepSeek | deepseek-v4-pro (peak rate) | $1.32 | $0.044 | $3.96 |
| DeepSeek | deepseek-flash, V4.1-Flash (peak rate) | $0.30 | $0.006 | $1.20 |

Three cautions apply to any such table. Tokens are not a common unit: Anthropic notes that the tokenizer used by its newer models produces approximately 30% more tokens for the same text than its earlier one, and vendors' tokenizers differ from each other. Cached-input figures are read prices and exclude Anthropic's write premium. And models differ in how many tokens they spend to finish the same task, so price per token and cost per completed task can rank models differently.

### The Trend

Within a capability tier, prices have fallen. Anthropic's own table shows the path: Claude Opus 4.1 at $15 input and $75 output, the Opus 4.5 through Opus 5 generation at $5 and $25, and Opus 5.5 at $4 and $20; Sonnet moved from $3 and $15 to $2 and $10, and Anthropic made the Sonnet 5 introductory price permanent rather than raising it in September 2026 as first announced. Three vendors now have a tier at $2 input and about $10 output. Below it, small and open-weight-derived endpoints sit one to two orders of magnitude cheaper: gpt-6-luna and deepseek-flash list at $0.10 to $0.30 per million input tokens.

The decline is not uniform. Anthropic and OpenAI each list a tier above their other models at $10 and $50. Google's page states that Gemini 3.8 Flash prices will double on 1 January 2027. Tokenizer changes can raise the token count for the same work. A fixed level of capability gets cheaper each generation, while the newest capability holds or raises its price.

### What Cheaper Capable Models Change for Agent Design

Cheap mid-tier models change which architectures are affordable. Running several attempts and selecting the best, adding a verifier pass, or delegating search to sub-agents on a small model in an [orchestrator-worker](https://metavert.io/orchestrator-worker-pattern) arrangement become routine line items. Model routing becomes a design decision: the same [harness](https://metavert.io/agent-harness) can send planning to an expensive model and bulk reading to one that costs a twentieth as much. Cheap input with deep cache discounts also weakens the cost argument for aggressive context pruning, leaving quality as the main reason to keep context small (see [context engineering](https://metavert.io/context-engineering)). The budgeting unit should be the task: tokens per successful outcome, including retries and reasoning, measured on the team's own workload.

### Build Versus Fine-Tune

Falling API prices raise the bar that a custom model must clear. A [fine-tuned](https://metavert.io/fine-tuning) small model used to win on cost almost automatically; when a capable hosted model lists at $0.30 per million input tokens, the saving from self-hosting may not cover GPU time, evaluation and maintenance unless volume is high and steady. The stronger remaining reasons to fine-tune an [open-weight model](https://metavert.io/open-weight-models) are not unit price: latency and deployment constraints, data that cannot leave an environment, behaviour that prompting cannot reliably produce, and independence from a vendor's deprecation schedule. Cheap frontier APIs also make fine-tuning easier, since a strong model can generate or grade training data for a small one (see [knowledge distillation](https://metavert.io/knowledge-distillation) and [reinforcement fine-tuning](https://metavert.io/reinforcement-fine-tuning)).

## Related Topics

- [Reasoning Models](https://metavert.io/reasoning-models) — Why hidden thinking tokens raise output cost
- [Test-Time Compute](https://metavert.io/test-time-compute) — Spending more tokens per task for better answers
- [Context Compaction](https://metavert.io/context-compaction) — Reducing input tokens, at a cost to cache hits
- [Context Windows](https://metavert.io/context-windows) — The limit that long-context pricing tiers attach to
- [Open-Weight Models](https://metavert.io/open-weight-models) — The self-hosted alternative to per-token pricing
- [DeepSeek](https://metavert.io/deepseek) — The low-price end of the hosted market
- [Fine-Tuning](https://metavert.io/fine-tuning) — The build option that API prices are weighed against
- [Orchestrator-Worker Pattern](https://metavert.io/orchestrator-worker-pattern) — Routing work across models of different cost
- [Small Language Models](https://metavert.io/small-language-models) — Where the cheapest tiers come from

## Further Reading

- [Pricing](https://platform.claude.com/docs/en/about-claude/pricing) — Anthropic, accessed October 2026
- [API pricing](https://developers.openai.com/api/docs/pricing) — OpenAI, accessed October 2026
- [Reasoning models guide](https://developers.openai.com/api/docs/guides/reasoning) — OpenAI, accessed October 2026
- [Gemini Developer API pricing](https://ai.google.dev/gemini-api/docs/pricing) — Google, October 2026
- [Models and pricing](https://api-docs.deepseek.com/quick_start/pricing) — DeepSeek, accessed October 2026
- [Context caching guide](https://api-docs.deepseek.com/guides/kv_cache) — DeepSeek, accessed October 2026
