OpenAIlegacy
GPT-4.1 pricing
$2.00 per million input tokens, $8.00 per million output. Read from OpenAI’s own documentation on August 3, 2026.
- Input
- $2.00 per 1M tokens
- Cached input
- $0.5000 25% of base
- Output
- $8.00 4.0× input
- Context window
- — not documented
- Token counting
- Exact o200k_base
- Price verified
- 2026-08-03 2 days ago
Source: OpenAI pricing documentation. Prices change without notice — verify before committing spend.
What GPT-4.1 costs on real work
Four workload shapes at 100,000 requests a month. The point of showing four is that the ranking between models changes depending on which one describes you.
| Workload | In | Out | Per request | Per month |
|---|---|---|---|---|
| ClassificationShort input, one-word answer. Input-dominated. | 500 | 50 | $0.001400 | $140.00 |
| Chat turnA system prompt plus a few turns of history. | 1,500 | 300 | $0.005400 | $540.00 |
| Document summaryA long document in, a paragraph out. | 20,000 | 800 | $0.0464 | $4,640.00 |
| Code generationOutput-heavy — where output pricing dominates. | 2,000 | 1,500 | $0.0160 | $1,600.00 |
Put your own numbers in the cost calculator, or measure a real prompt first in the token counter. If your requests share a stable prefix, the cached rate applies to most of your input — check the structure in the cache checker.
Counting tokens for GPT-4.1
GPT-4.1 uses the o200k_base encoding, which OpenAI publishes. That means a count taken before you send is exact — the same number the API bills you for.
The token counter runs that encoder in your browser, so nothing is uploaded and the figure needs no caveat.
Other OpenAI models
The tier question: is a cheaper model in the same family enough for your task?
| Model | Input | Output | Context | Chat turn |
|---|---|---|---|---|
| GPT-4.1 — this page | $2.00 | $8.00 | — | $0.005400 |
| GPT-5.6 Sol | $5.00 | $30.00 | 1.05M | $0.0165 |
| GPT-5.6 Terra | $2.00 | $12.00 | 1.05M | $0.006600 |
| GPT-5.6 Luna | $0.2000 | $1.20 | 1.05M | $0.000660 |
| GPT-5.5 | $5.00 | $30.00 | — | $0.0165 |
| GPT-5.4 | $2.50 | $15.00 | — | $0.008250 |
| GPT-5.4 mini | $0.7500 | $4.50 | — | $0.002475 |
Alternatives from other providers
Models priced nearest to GPT-4.1, not the cheapest on the market — those are the ones actually worth evaluating against it.
Frequently asked questions
- How much does GPT-4.1 cost?
- $2.00 per million input tokens and $8.00 per million output tokens, with cached input at $0.5000 per million. On a typical chat turn of 1,500 input and 300 output tokens that is $0.005400 per request, or $540.00 per month at 100,000 requests. Read from OpenAI's own documentation on August 3, 2026.
- Can I count GPT-4.1 tokens exactly?
- Yes. GPT-4.1 uses the o200k_base encoding, which is published and runs in a browser — so a count taken before you send is the number you will be billed for.
- Why is output more expensive than input on GPT-4.1?
- Output costs 4.0 times input here. Input is processed in a single parallel pass, while output is generated one token at a time with a full pass over the model for each. That is why a model that answers concisely can be cheaper in production than one with a lower headline rate.