Claude Sonnet 5 pricing
$2.00 per million input tokens, $10.00 per million output. 1M token context window. Read from Anthropic’s own documentation on August 3, 2026.
- Input
- $2.00 per 1M tokens
- Cached input
- $0.2000 10% of base
- Output
- $10.00 5.0× input
- Context window
- 1M 128K max output
- Token counting
- Estimate o200k_base
- Price verified
- 2026-08-03 2 days ago
Introductory pricing through 31 Aug 2026. From 1 Sep 2026: $3 input / $15 output per 1M tokens.
Source: Anthropic pricing documentation. Prices change without notice — verify before committing spend.
What Claude Sonnet 5 costs on real work
Four workload shapes at 100,000 requests a month. The point of showing four is that the ranking between models changes depending on which one describes you.
| Workload | In | Out | Per request | Per month |
|---|---|---|---|---|
| ClassificationShort input, one-word answer. Input-dominated. | 500 | 50 | $0.001500 | $150.00 |
| Chat turnA system prompt plus a few turns of history. | 1,500 | 300 | $0.006000 | $600.00 |
| Document summaryA long document in, a paragraph out. | 20,000 | 800 | $0.0480 | $4,800.00 |
| Code generationOutput-heavy — where output pricing dominates. | 2,000 | 1,500 | $0.0190 | $1,900.00 |
Put your own numbers in the cost calculator, or measure a real prompt first in the token counter. If your requests share a stable prefix, the cached rate applies to most of your input — check the structure in the cache checker.
Counting tokens for Claude Sonnet 5
Anthropic does not publish a tokenizer that runs in a browser, so any pre-flight count for Claude Sonnet 5 is an estimate rather than a measurement.
Anthropic does not publish a client-side tokenizer. Counted with o200k_base and scaled ~1.18x, the commonly reported gap between tiktoken and Claude's pre-4.7 tokenizer.
Treat it as accurate to within roughly ten to twenty percent. That is fine for budgeting and wrong for sizing a prompt right at a context window boundary — where precision matters, use Anthropic’s own token counting endpoint from your backend. The methodology page sets out every scaling factor used here.
Other Anthropic models
The tier question: is a cheaper model in the same family enough for your task?
| Model | Input | Output | Context | Chat turn |
|---|---|---|---|---|
| Claude Sonnet 5 — this page | $2.00 | $10.00 | 1M | $0.006000 |
| Claude Fable 5 | $10.00 | $50.00 | 1M | $0.0300 |
| Claude Opus 5 | $5.00 | $25.00 | 1M | $0.0150 |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | $0.0150 |
| Claude Opus 4.6 | $5.00 | $25.00 | 1M | $0.0150 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | $0.009000 |
| Claude Sonnet 4.5 | $3.00 | $15.00 | 200K | $0.009000 |
Alternatives from other providers
Models priced nearest to Claude Sonnet 5, not the cheapest on the market — those are the ones actually worth evaluating against it.
- OpenAIGPT-5.6 Terra$2.00 in · $12.00 out↑ 10% on a chat turn
- OpenAIGPT-5.1$1.25 in · $10.00 out↓ 19% on a chat turn
- OpenAIGPT-5$1.25 in · $10.00 out↓ 19% on a chat turn
- GoogleGemini 3.5 Flash$1.50 in · $9.00 out↓ 17% on a chat turn
Side-by-side comparisons: Claude Sonnet 5 vs GPT-5.4 · Claude Sonnet 5 vs Gemini 3.5 Flash
Frequently asked questions
- How much does Claude Sonnet 5 cost?
- $2.00 per million input tokens and $10.00 per million output tokens, with cached input at $0.2000 per million. On a typical chat turn of 1,500 input and 300 output tokens that is $0.006000 per request, or $600.00 per month at 100,000 requests. Read from Anthropic's own documentation on August 3, 2026.
- Can I count Claude Sonnet 5 tokens exactly?
- No. Anthropic does not publish a tokenizer that runs in a browser, so any pre-flight count for Claude Sonnet 5 is an estimate. Anthropic does not publish a client-side tokenizer. Counted with o200k_base and scaled ~1.18x, the commonly reported gap between tiktoken and Claude's pre-4.7 tokenizer. Treat it as accurate to within roughly ten to twenty percent and never as the basis for sizing a prompt right at a context window boundary.
- What is the context window of Claude Sonnet 5?
- 1,000,000 tokens, with a maximum of 128,000 output tokens in a single response. That budget covers everything in the request — system prompt, conversation history, tool definitions, documents — plus the response itself, not just your input.
- Why is output more expensive than input on Claude Sonnet 5?
- Output costs 5.0 times input here. Input is processed in a single parallel pass, while output is generated one token at a time with a full pass over the model for each. That is why a model that answers concisely can be cheaper in production than one with a lower headline rate.