TokenPad

Mistral AI

Mistral Small 4 pricing

$0.1500 per million input tokens, $0.6000 per million output. 256K token context window. Read from Mistral AI’s own documentation on August 5, 2026.

Input
$0.1500
per 1M tokens
Cached input
not published
Output
$0.6000
4.0× input
Context window
256K
verified
Token counting
Estimate
o200k_base
Price verified
2026-08-05
today

Source: Mistral AI pricing documentation. Prices change without notice — verify before committing spend.

What Mistral Small 4 costs on real work

Four workload shapes at 100,000 requests a month. The point of showing four is that the ranking between models changes depending on which one describes you.

Mistral Small 4 cost by workload shape
WorkloadInOutPer requestPer month
ClassificationShort input, one-word answer. Input-dominated.50050$0.000105$10.50
Chat turnA system prompt plus a few turns of history.1,500300$0.000405$40.50
Document summaryA long document in, a paragraph out.20,000800$0.003480$348.00
Code generationOutput-heavy — where output pricing dominates.2,0001,500$0.001200$120.00

Put your own numbers in the cost calculator, or measure a real prompt first in the token counter.

Counting tokens for Mistral Small 4

Mistral AI does not publish a tokenizer that runs in a browser, so any pre-flight count for Mistral Small 4 is an estimate rather than a measurement.

Mistral publishes its Tekken tokenizer only as a Python package, with no browser build. Counted with o200k_base and scaled; Tekken generally produces more tokens than o200k_base on the same English text.

Treat it as accurate to within roughly ten to twenty percent. That is fine for budgeting and wrong for sizing a prompt right at a context window boundary — where precision matters, use Mistral AI’s own token counting endpoint from your backend. The methodology page sets out every scaling factor used here.

Other Mistral AI models

The tier question: is a cheaper model in the same family enough for your task?

Other Mistral AI models compared with Mistral Small 4
ModelInputOutputContextChat turn
Mistral Small 4 — this page$0.1500$0.6000256K$0.000405
Mistral Medium 3.5$1.50$7.50256K$0.004500
Mistral Large 3$0.5000$1.50256K$0.001200
Magistral Medium$2.00$5.00256K$0.004500
Devstral 2$0.4000$2.00256K$0.001200
Ministral 3 8B$0.1500$0.1500256K$0.000270

Alternatives from other providers

Models priced nearest to Mistral Small 4, not the cheapest on the market — those are the ones actually worth evaluating against it.

Frequently asked questions

How much does Mistral Small 4 cost?
$0.1500 per million input tokens and $0.6000 per million output tokens. On a typical chat turn of 1,500 input and 300 output tokens that is $0.000405 per request, or $40.50 per month at 100,000 requests. Read from Mistral AI's own documentation on August 5, 2026.
Can I count Mistral Small 4 tokens exactly?
No. Mistral AI does not publish a tokenizer that runs in a browser, so any pre-flight count for Mistral Small 4 is an estimate. Mistral publishes its Tekken tokenizer only as a Python package, with no browser build. Counted with o200k_base and scaled; Tekken generally produces more tokens than o200k_base on the same English text. Treat it as accurate to within roughly ten to twenty percent and never as the basis for sizing a prompt right at a context window boundary.
What is the context window of Mistral Small 4?
256,000 tokens. That budget covers everything in the request — system prompt, conversation history, tool definitions, documents — plus the response itself, not just your input.
Why is output more expensive than input on Mistral Small 4?
Output costs 4.0 times input here. Input is processed in a single parallel pass, while output is generated one token at a time with a full pass over the model for each. That is why a model that answers concisely can be cheaper in production than one with a lower headline rate.