Gemini 3.5 Flash pricing
$1.50 per million input tokens, $9.00 per million output. Read from Google’s own documentation on August 3, 2026.
- Input
- $1.50 per 1M tokens
- Cached input
- — not published
- Output
- $9.00 6.0× input
- Context window
- — not documented
- Token counting
- Estimate o200k_base
- Price verified
- 2026-08-03 2 days ago
Source: Google pricing documentation. Prices change without notice — verify before committing spend.
What Gemini 3.5 Flash costs on real work
Four workload shapes at 100,000 requests a month. The point of showing four is that the ranking between models changes depending on which one describes you.
| Workload | In | Out | Per request | Per month |
|---|---|---|---|---|
| ClassificationShort input, one-word answer. Input-dominated. | 500 | 50 | $0.001200 | $120.00 |
| Chat turnA system prompt plus a few turns of history. | 1,500 | 300 | $0.004950 | $495.00 |
| Document summaryA long document in, a paragraph out. | 20,000 | 800 | $0.0372 | $3,720.00 |
| Code generationOutput-heavy — where output pricing dominates. | 2,000 | 1,500 | $0.0165 | $1,650.00 |
Put your own numbers in the cost calculator, or measure a real prompt first in the token counter.
Counting tokens for Gemini 3.5 Flash
Google does not publish a tokenizer that runs in a browser, so any pre-flight count for Gemini 3.5 Flash is an estimate rather than a measurement.
Gemini uses a SentencePiece tokenizer that has no browser build. Counted with o200k_base at parity; no published conversion factor exists, so treat this as an order-of-magnitude figure.
Treat it as accurate to within roughly ten to twenty percent. That is fine for budgeting and wrong for sizing a prompt right at a context window boundary — where precision matters, use Google’s own token counting endpoint from your backend. The methodology page sets out every scaling factor used here.
Other Google models
The tier question: is a cheaper model in the same family enough for your task?
| Model | Input | Output | Context | Chat turn |
|---|---|---|---|---|
| Gemini 3.5 Flash — this page | $1.50 | $9.00 | — | $0.004950 |
| Gemini 3.6 Flash | $1.50 | $7.50 | — | $0.004500 |
| Gemini 3.5 Flash-Lite | $0.3000 | $2.50 | — | $0.001200 |
| Gemini 2.5 Flash | $0.3000 | $2.50 | 1M | $0.001200 |
| Gemini 2.5 Flash-Lite | $0.1000 | $0.4000 | — | $0.000270 |
Alternatives from other providers
Models priced nearest to Gemini 3.5 Flash, not the cheapest on the market — those are the ones actually worth evaluating against it.
- OpenAIGPT-5.1$1.25 in · $10.00 out↓ 2% on a chat turn
- OpenAIGPT-5$1.25 in · $10.00 out↓ 2% on a chat turn
- xAIGrok 4.5$2.00 in · $6.00 out↓ 3% on a chat turn
- Mistral AIMistral Medium 3.5$1.50 in · $7.50 out↓ 9% on a chat turn
Side-by-side comparisons: Gemini 3.5 Flash vs GPT-5 mini · Claude Haiku 4.5 vs Gemini 3.5 Flash · Claude Sonnet 5 vs Gemini 3.5 Flash
Frequently asked questions
- How much does Gemini 3.5 Flash cost?
- $1.50 per million input tokens and $9.00 per million output tokens. On a typical chat turn of 1,500 input and 300 output tokens that is $0.004950 per request, or $495.00 per month at 100,000 requests. Read from Google's own documentation on August 3, 2026.
- Can I count Gemini 3.5 Flash tokens exactly?
- No. Google does not publish a tokenizer that runs in a browser, so any pre-flight count for Gemini 3.5 Flash is an estimate. Gemini uses a SentencePiece tokenizer that has no browser build. Counted with o200k_base at parity; no published conversion factor exists, so treat this as an order-of-magnitude figure. Treat it as accurate to within roughly ten to twenty percent and never as the basis for sizing a prompt right at a context window boundary.
- Why is output more expensive than input on Gemini 3.5 Flash?
- Output costs 6.0 times input here. Input is processed in a single parallel pass, while output is generated one token at a time with a full pass over the model for each. That is why a model that answers concisely can be cheaper in production than one with a lower headline rate.