TokenPad
Cost

LLM Model Finder

Filter every model by budget, context and provider in one pass.

Ranked by the cost of one representative request at your ratio — 4 input tokens for every output token. That ranking is not the same as sorting by input price, and for output-heavy work it frequently inverts it.

32 of 42 models match · 10 filtered out

ModelCountInput $/1MOutput $/1MContextSample request
GPT-5 nanoOpenAIExact$0.0500$0.4000$0.000600
Ministral 3 8BMistral AIEstimate$0.1500$0.1500256K$0.000750
DeepSeek V4 FlashDeepSeekEstimate$0.1400$0.28001M$0.000840
Mistral Small 4Mistral AIEstimate$0.1500$0.6000256K$0.001200
GPT-5.6 LunaOpenAIExact$0.2000$1.201.05M$0.002000
GPT-5.4 nanoOpenAIExact$0.2000$1.25$0.002050
DeepSeek V4 ProDeepSeekEstimate$0.4350$0.87001M$0.002610
GPT-5 miniOpenAIExact$0.2500$2.00$0.003000
Mistral Large 3Mistral AIEstimate$0.5000$1.50256K$0.003500
Devstral 2Mistral AIEstimate$0.4000$2.00256K$0.003600
Gemini 3.5 Flash-LiteGoogleEstimate$0.3000$2.50$0.003700
Grok Build 0.1xAIEstimate$1.00$2.00256K$0.006000
GPT-5.4 miniOpenAIExact$0.7500$4.50$0.007500
Grok 4.3xAIEstimate$1.25$2.501M$0.007500
Claude Haiku 4.5AnthropicEstimate$1.00$5.00200K$0.009000
Magistral MediumMistral AIEstimate$2.00$5.00256K$0.0130
Gemini 3.6 FlashGoogleEstimate$1.50$7.50$0.0135
Mistral Medium 3.5Mistral AIEstimate$1.50$7.50256K$0.0135
Grok 4.5xAIEstimate$2.00$6.00500K$0.0140
GPT-5.1OpenAIExact$1.25$10.00$0.0150
GPT-5OpenAIExact$1.25$10.00$0.0150
Gemini 3.5 FlashGoogleEstimate$1.50$9.00$0.0150
Claude Sonnet 5AnthropicEstimate$2.00$10.001M$0.0180
GPT-5.6 TerraOpenAIExact$2.00$12.001.05M$0.0200
GPT-5.4OpenAIExact$2.50$15.00$0.0250
Claude Sonnet 4.6AnthropicEstimate$3.00$15.001M$0.0270
Claude Opus 5AnthropicEstimate$5.00$25.001M$0.0450
Claude Opus 4.8AnthropicEstimate$5.00$25.001M$0.0450
Claude Opus 4.6AnthropicEstimate$5.00$25.001M$0.0450
GPT-5.6 SolOpenAIExact$5.00$30.001.05M$0.0500
GPT-5.5OpenAIExact$5.00$30.00$0.0500
Claude Fable 5AnthropicEstimate$10.00$50.001M$0.0900

This narrows the field; it does not pick for you. Capability differences between the survivors will not appear in any price table. Take the top three or four into the cost calculator with your real volume, then evaluate them on your own task.

What this tool tells you

Thirty-odd models, four providers, three prices each. This narrows that field to the handful worth evaluating: set a budget, a context requirement and your input-to-output ratio, and the survivors come back ranked by what a representative request actually costs you.

It is a shortlisting tool, not a recommendation engine. Capability differences between the survivors will not appear in any price table, and no amount of filtering substitutes for running your own task against three candidates.

The ratio slider is the important control

Everything else is a filter. The ratio is what changes the ranking, and it is the number most people have never measured for their own workload.

  • High ratio (10:1 and above) — summarisation, classification, extraction, retrieval-augmented answering. Input rate dominates and output pricing barely registers.
  • Low ratio (2:1 and below) — content generation, code writing, translation. Output rate dominates completely, and the ranking from the high-ratio case frequently inverts.

Move the slider from 20 to 1 and watch the table reorder. That reordering is the entire reason a generic "cheapest models" list is not useful to you. Get your real ratio from production logs, or measure a representative request in the token counter.

What each filter does

Price caps

Hard ceilings on the published input and output rates. Useful for eliminating the flagship tier when you already know the task does not need it.

Minimum context

Filters on the verified window. Note that this excludes every model whose window we have not been able to confirm against the provider's own documentation — an omission rather than an invented figure, for reasons set out on the methodology page. If a model you expected is missing, that is why.

Exact token counts only

Restricts to models whose tokenizer can run in a browser, so pre-flight counts are exact rather than estimated. Worth switching on if you need to size prompts close to a context limit, or if your cost forecasting has to be defensible to someone else.

Hide legacy models

On by default. Superseded models often look attractive on price and are a poor place to start a new build, since they get deprecated on the provider's timeline rather than yours.

After the shortlist

  1. Take the top three or four into the cost calculator with your real monthly volume and cacheable share. A model that wins on a sample request can lose once caching is applied.
  2. Confirm your actual context fits in the context window calculator.
  3. Check your throughput against the quota in the rate limit calculator — a model you cannot get enough capacity on is not a candidate regardless of price.
  4. Then evaluate on your own task. Price differences are irrelevant among models that fail it.

One further note worth internalising: optimising the workload usually beats switching models. Before committing to a migration, read how to reduce LLM API costs — the first three items there routinely return more than any provider change.

Frequently asked questions

Why sort by blended cost rather than input price?
Because almost nobody has an input-only workload. A model can be cheapest on input and dearest overall once its output rate is applied at your real ratio. The sort here uses the input-to-output mix you set, which is the only ranking that describes your bill rather than someone else’s.
Why do some models have no context window listed?
Because the provider does not document it plainly enough to publish. Those models are excluded from context filtering rather than assigned a plausible figure. An omission is honest; an invented number in a tool people use to size production prompts is not.
Should I always pick the cheapest match?
No. This narrows a field of thirty-odd models to a shortlist of three or four worth evaluating. Capability differences between them will not appear in any price table, and a weaker model that needs two attempts or a human correction costs more than its rate suggests.

More cost tools