LLM API pricing, with sources
42 models from 6 providers. Every price below was read from the provider’s own documentation — each row carries the date it was checked and a link to the page it came from, so you can verify any figure in one click instead of trusting us. Last full review: August 5, 2026.
Find a model by what you actually need
42 of 42 models match. Cheapest on a 3:1 input-to-output mix is GPT-5 nano at $0.138 per million blended.
- GPT-5 nanoOpenAIin$0.050out$0.400ctx—exact
- Ministral 3 8BMistral AIin$0.150out$0.150ctx256Kestimate
- Gemini 2.5 Flash-LiteGooglein$0.100out$0.400ctx—estimate
- DeepSeek V4 FlashDeepSeekin$0.140out$0.280ctx1Mestimate
- GPT-4o miniOpenAIin$0.150out$0.600ctx—exact
- Mistral Small 4Mistral AIin$0.150out$0.600ctx256Kestimate
- GPT-5.6 LunaOpenAIin$0.200out$1.20ctx1.1Mexact
- GPT-5.4 nanoOpenAIin$0.200out$1.25ctx—exact
- DeepSeek V4 ProDeepSeekin$0.435out$0.870ctx1Mestimate
- GPT-5 miniOpenAIin$0.250out$2.00ctx—exact
- GPT-4.1 miniOpenAIin$0.400out$1.60ctx—exact
- GPT-3.5 TurboOpenAIin$0.500out$1.50ctx—exact
- Mistral Large 3Mistral AIin$0.500out$1.50ctx256Kestimate
- Devstral 2Mistral AIin$0.400out$2.00ctx256Kestimate
- Gemini 3.5 Flash-LiteGooglein$0.300out$2.50ctx—estimate
- Gemini 2.5 FlashGooglein$0.300out$2.50ctx1Mestimate
- Grok Build 0.1xAIin$1.00out$2.00ctx256Kestimate
- Grok 4.3xAIin$1.25out$2.50ctx1Mestimate
- GPT-5.4 miniOpenAIin$0.750out$4.50ctx—exact
- o4-miniOpenAIin$1.10out$4.40ctx—exact
- Claude Haiku 4.5Anthropicin$1.00out$5.00ctx200Kestimate
- Magistral MediumMistral AIin$2.00out$5.00ctx256Kestimate
- Gemini 3.6 FlashGooglein$1.50out$7.50ctx—estimate
- Grok 4.5xAIin$2.00out$6.00ctx500Kestimate
- Mistral Medium 3.5Mistral AIin$1.50out$7.50ctx256Kestimate
- Gemini 3.5 FlashGooglein$1.50out$9.00ctx—estimate
- GPT-5.1OpenAIin$1.25out$10.00ctx—exact
- GPT-5OpenAIin$1.25out$10.00ctx—exact
- GPT-4.1OpenAIin$2.00out$8.00ctx—exact
- o3OpenAIin$2.00out$8.00ctx—exact
- Claude Sonnet 5Anthropicin$2.00out$10.00ctx1Mestimate
- GPT-4oOpenAIin$2.50out$10.00ctx—exact
- GPT-5.6 TerraOpenAIin$2.00out$12.00ctx1.1Mexact
- GPT-5.4OpenAIin$2.50out$15.00ctx—exact
- Claude Sonnet 4.6Anthropicin$3.00out$15.00ctx1Mestimate
- Claude Sonnet 4.5Anthropicin$3.00out$15.00ctx200Kestimate
- Claude Opus 5Anthropicin$5.00out$25.00ctx1Mestimate
- Claude Opus 4.8Anthropicin$5.00out$25.00ctx1Mestimate
- Claude Opus 4.6Anthropicin$5.00out$25.00ctx1Mestimate
- GPT-5.6 SolOpenAIin$5.00out$30.00ctx1.1Mexact
- GPT-5.5OpenAIin$5.00out$30.00ctx—exact
- Claude Fable 5Anthropicin$10.00out$50.00ctx1Mestimate
OpenAI
18 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $0.5000 | $30.00 | 1.05M | Exact | 2026-08-03 |
| GPT-5.6 Terra | $2.00 | $0.2000 | $12.00 | 1.05M | Exact | 2026-08-03 |
| GPT-5.6 Luna | $0.2000 | $0.0200 | $1.20 | 1.05M | Exact | 2026-08-03 |
| GPT-5.5 | $5.00 | $0.5000 | $30.00 | — | Exact | 2026-08-03 |
| GPT-5.4 | $2.50 | $0.2500 | $15.00 | — | Exact | 2026-08-03 |
| GPT-5.4 mini | $0.7500 | $0.0750 | $4.50 | — | Exact | 2026-08-03 |
| GPT-5.4 nano | $0.2000 | $0.0200 | $1.25 | — | Exact | 2026-08-03 |
| GPT-5.1 | $1.25 | $0.1250 | $10.00 | — | Exact | 2026-08-03 |
| GPT-5 | $1.25 | $0.1250 | $10.00 | — | Exact | 2026-08-03 |
| GPT-5 mini | $0.2500 | $0.0250 | $2.00 | — | Exact | 2026-08-03 |
| GPT-5 nano | $0.0500 | $0.005000 | $0.4000 | — | Exact | 2026-08-03 |
| GPT-4.1 legacy | $2.00 | $0.5000 | $8.00 | — | Exact | 2026-08-03 |
| GPT-4.1 mini legacy | $0.4000 | $0.1000 | $1.60 | — | Exact | 2026-08-03 |
| GPT-4o legacy | $2.50 | $1.25 | $10.00 | — | Exact | 2026-08-03 |
| GPT-4o mini legacy | $0.1500 | $0.0750 | $0.6000 | — | Exact | 2026-08-03 |
| o3 legacy | $2.00 | $0.5000 | $8.00 | — | Exact | 2026-08-03 |
| o4-mini legacy | $1.10 | $0.2750 | $4.40 | — | Exact | 2026-08-03 |
| GPT-3.5 Turbo legacy | $0.5000 | — | $1.50 | — | Exact | 2026-08-03 |
Source: OpenAI pricing documentation
Anthropic
8 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $1.00 | $50.00 | 1M | Estimate | 2026-08-03 |
| Claude Opus 5 | $5.00 | $0.5000 | $25.00 | 1M | Estimate | 2026-08-03 |
| Claude Opus 4.8 | $5.00 | $0.5000 | $25.00 | 1M | Estimate | 2026-08-03 |
| Claude Opus 4.6 | $5.00 | $0.5000 | $25.00 | 1M | Estimate | 2026-08-03 |
| Claude Sonnet 5Introductory pricing through 31 Aug 2026. From 1 Sep 2026: $3 input / $15 output per 1M tokens. | $2.00 | $0.2000 | $10.00 | 1M | Estimate | 2026-08-03 |
| Claude Sonnet 4.6 | $3.00 | $0.3000 | $15.00 | 1M | Estimate | 2026-08-03 |
| Claude Sonnet 4.5 legacy | $3.00 | $0.3000 | $15.00 | 200K | Estimate | 2026-08-03 |
| Claude Haiku 4.5 | $1.00 | $0.1000 | $5.00 | 200K | Estimate | 2026-08-03 |
Source: Anthropic pricing documentation
5 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| Gemini 3.6 Flash | $1.50 | — | $7.50 | — | Estimate | 2026-08-03 |
| Gemini 3.5 Flash | $1.50 | — | $9.00 | — | Estimate | 2026-08-03 |
| Gemini 3.5 Flash-Lite | $0.3000 | — | $2.50 | — | Estimate | 2026-08-03 |
| Gemini 2.5 Flash legacy | $0.3000 | — | $2.50 | 1M | Estimate | 2026-08-03 |
| Gemini 2.5 Flash-Lite legacy | $0.1000 | — | $0.4000 | — | Estimate | 2026-08-03 |
Source: Google pricing documentation
DeepSeek
2 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | $0.1400 | $0.002800 | $0.2800 | 1M | Estimate | 2026-08-03 |
| DeepSeek V4 Pro | $0.4350 | $0.003625 | $0.8700 | 1M | Estimate | 2026-08-03 |
Source: DeepSeek pricing documentation
xAI
3 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| Grok 4.5Rates shown are for prompts under 200K tokens. Above that threshold xAI charges double on every line: $4.00 input, $0.60 cached, $12.00 output. | $2.00 | $0.3000 | $6.00 | 500K | Estimate | 2026-08-05 |
| Grok 4.3Rates shown are for prompts under 200K tokens. Above that threshold every rate doubles: $2.50 input, $0.40 cached, $5.00 output. | $1.25 | $0.2000 | $2.50 | 1M | Estimate | 2026-08-05 |
| Grok Build 0.1Rates shown are for prompts under 200K tokens; above that they double to $2.00 input and $4.00 output. | $1.00 | $0.2000 | $2.00 | 256K | Estimate | 2026-08-05 |
Source: xAI pricing documentation
Mistral AI
6 models · pricing quirks and token counting →
| Model | Input | Cached in | Output | Context | Token count | Verified |
|---|---|---|---|---|---|---|
| Mistral Medium 3.5Mistral advertises a 90% discount on cached input tokens but does not publish a per-model cached rate, so no cached figure is shown here rather than an inferred one. | $1.50 | — | $7.50 | 256K | Estimate | 2026-08-05 |
| Mistral Large 3Open-weight model, so self-hosting is an alternative to these API rates. Cached input is discounted but not published per model. | $0.5000 | — | $1.50 | 256K | Estimate | 2026-08-05 |
| Mistral Small 4 | $0.1500 | — | $0.6000 | 256K | Estimate | 2026-08-05 |
| Magistral MediumReasoning model. Reasoning tokens are billed at the output rate and do not appear in the visible answer. | $2.00 | — | $5.00 | 256K | Estimate | 2026-08-05 |
| Devstral 2Aimed at coding and agentic work. Code tokenizes denser than prose, so token counts for the same character count run higher here than on text workloads. | $0.4000 | — | $2.00 | 256K | Estimate | 2026-08-05 |
| Ministral 3 8BInput and output are priced identically, which is unusual and makes output-heavy work disproportionately cheap here relative to models that charge a premium on output. | $0.1500 | — | $0.1500 | 256K | Estimate | 2026-08-05 |
Source: Mistral AI pricing documentation
How to use this table
All figures are US dollars per million tokens, at standard on-demand rates. Batch processing, enterprise agreements and regional premiums are not reflected. Cached input is what you pay for prompt content the provider has already processed and retained — usually about a tenth of the base input rate, and the largest lever available on a repetitive workload.
The token count column is the one most price tables omit. A rate per million tokens is only meaningful if you know how many tokens your text becomes, and that depends on a tokenizer that not every provider publishes. Where we can run the real encoder, the column says exact; where we cannot, it says estimate. The methodology page explains what sits behind each label.
Put your own numbers against these rates in the cost calculator, or measure a real prompt first in the token counter.
If you are choosing between two specific models rather than surveying the field, the head-to-head comparisons price both across four workload shapes at a million requests a month — which is where a ranking taken from the headline rate sometimes reverses.