YFarmX logoYFarmX

Cost estimator

What your workload costs on every measured model

A bill for an AI model is three numbers added up: the tokens it reads, the tokens it writes back, and the tokens it spends thinking before it answers. This tool works out all three for 308 models at once. Pick the workload closest to yours or paste a sample of your own prompts, set how many requests a month you send, and the table below redraws, cheapest first.

The input side is measured against real bills. Models that share a tokenizer count any text the same way, and the fingerprint dataset says which ones do, so we billed 24 real documents (articles, code, JSON, CSV, HTML and encyclopedia text in nine writing systems) on one model from each of 50 tokenizer families and fitted each family's rate for every kind of text. Tested on 18 other documents billed on 12 models, the input count landed 5.8% from the bill at the median and within 10% on 80% of readings.

The thinking side is measured: between 5 October 2026 and 8 October 2026 we sent eight tasks to 44 reasoning models, three times each at their default effort and once each at low and high, and read the reasoning tokens from the bill. The table shows the median and the 90th percentile for the task closest to your workload and how many runs they come from; with three runs, the 90th percentile is the longest of the three.

Showing Customer support chat: 1,200 characters a prompt, 150 output tokens, 100,000 requests a month, effort as called. Prices as listed on OpenRouter on 8 October 2026.

ModelInput tokensThinking tokens, p50 to p90Per requestPer month, p50 to p90Note
Mistral: Mistral NemoMistral AI · Mistral vocabulary364noneUS$0.000011US$1
Sao10K: Llama 3 8B LunarisSao10k · Llama3 vocabulary350noneUS$0.000022US$2
IBM: Granite 4.0 MicroIBM · OpenAI cl100k vocabulary369noneUS$0.000023US$2
Meta: Llama 3.1 8B InstructMeta · Llama3 vocabulary349noneUS$0.000029US$3
Google: Gemma 3 4BGoogle · Gemini vocabulary358noneUS$0.000033US$3
Amazon: Nova Micro 1.0Amazon · Nova vocabulary361noneUS$0.000034US$3
Inference.net: Schematron V2 TurboInference Net · IBM vocabulary377noneUS$0.000034US$3
Cohere: Command R7B (12-2024)Cohere359noneUS$0.000036US$4
Mistral: Mistral Small 3Mistral AI524noneUS$0.000038US$4
Meta: Llama 3.2 1B InstructMeta · Llama3 vocabulary350noneUS$0.000040US$4
Google: Gemma 3 12BGoogle · Gemini vocabulary358noneUS$0.000040US$4
Tencent: Hy-MT2-1.8BTencent347noneUS$0.000042US$4
Microsoft: Phi 4Microsoft · OpenAI cl100k vocabulary346noneUS$0.000045US$5
Reka EdgeRekaai344noneUS$0.000049US$5
MythoMax 13BGryphe · Llama2 vocabulary426noneUS$0.000051US$5
Mistral: Ministral 3 3B 2512Mistral AI · Mistral vocabulary364noneUS$0.000051US$5
Inference.net: Schematron V2 SmallInference Net · Llama3 vocabulary374noneUS$0.000053US$5
Amazon: Nova Lite 1.0Amazon · Nova vocabulary361noneUS$0.000058US$6
Meta: Llama 3.2 3B InstructMeta · Llama3 vocabulary350noneUS$0.000067US$7
Qwen: Qwen3 Coder 30B A3B InstructAlibaba (Qwen) · Qwen vocabulary364noneUS$0.000067US$7
ByteDance: UI-TARS 7B Bytedance · Qwen vocabulary375noneUS$0.000067US$7
Qwen: Qwen2.5 7B InstructAlibaba (Qwen) · Qwen vocabulary385noneUS$0.000069US$7
Tencent: Hy-MT2-30B-A3BTencent356noneUS$0.000071US$7
Tencent: Hy-MT2-7BTencent359noneUS$0.000071US$7
Mistral: Mistral Small 3.2 24BMistral AI · Mistral vocabulary364noneUS$0.000072US$7
Mistral: Ministral 3 8B 2512Mistral AI · Mistral vocabulary364noneUS$0.000077US$8
Meta: Llama 4 ScoutMeta · Llama4 vocabulary343noneUS$0.000079US$8
Mistral: Voxtral Small 24B 2507Mistral AI · Mistral vocabulary364noneUS$0.000081US$8
Qwen: Qwen3 30B A3B Instruct 2507Alibaba (Qwen) · Qwen vocabulary364noneUS$0.000081US$8
Meta: Llama 3.3 70B InstructMeta · Llama3 vocabulary349noneUS$0.000083US$8
OpenAI: GPT-4.1 NanoOpenAI · GPT vocabulary342noneUS$0.000094US$9
Google: Gemma 3 27BGoogle · Gemini vocabulary358noneUS$0.000096US$10
Qwen: Qwen3 VL 32B InstructAlibaba (Qwen) · Qwen vocabulary364noneUS$0.000100US$10
Mistral: Ministral 3 14B 2512Mistral AI · Mistral vocabulary364noneUS$0.000103US$10
Qwen: Qwen3 VL 8B InstructAlibaba (Qwen) · Qwen vocabulary364noneUS$0.000111US$11
Qwen: Qwen3 235B A22B Instruct 2507Alibaba (Qwen) · Qwen vocabulary364noneUS$0.000115US$12
OpenAI: GPT-6 LunaOpenAI · GPT vocabulary34113 to 213 runsUS$0.000116US$12 to US$12
Meta: Llama Guard 4 12BMeta538noneUS$0.000124US$12
Anthropic: Claude Haiku 5.5Anthropic · Claude vocabulary51903 runsUS$0.000127US$13
OpenAI: GPT-4o-miniOpenAI · GPT vocabulary342noneUS$0.000141US$14

Showing the 40 cheapest of 308. Use the filter above to find a model, or .

The monthly range multiplies one request by the requests a month: every request at the median, then every request at the 90th percentile. A real month mixes short and long thinking, so its bill usually falls between the two.

Why thinking is a range

A reasoning model decides for itself how long to think, and the decision depends on the task and on the effort setting, so one request can cost ten times the next. The way to budget for that is to measure the distribution rather than guess a point. The think bench sends eight tasks (a code fix, a maths word problem, a JSON extraction, a summary, a classification, a question over a long passage, a tool-call plan and a translation) to each reasoning model at its default effort three times and at low and high once, and reads reasoning_tokens from the usage block of every answer. The table prints the median and the 90th percentile for the task type closest to your workload, with the number of runs behind them. Three runs per task will not catch the rare very long answer, so treat the 90th percentile as a planning figure, and cap reasoning in your own calls at the level you budget for.

A model whose listing accepts a reasoning setting but has not yet been through the bench is marked unmeasured, its figure is input and output only, and it is ranked after every model with a complete estimate, so a part of a bill never sits among whole ones. Where a model reasons at its default without being asked, the default row carries that; where it reasons only when asked, the default row reads zero and the low and high rows show what asking costs.

How the figures are made

Part of the billWhere the number comes from
Input tokensYour prompt is split the way a tokenizer splits it before it encodes anything: English, code and data into words, runs of digits, punctuation and whitespace, and Chinese, Japanese, Korean and other writing systems into characters. Each piece is priced at the model's own rate for its kind: prose, source code, structured data, numbers, whitespace, emoji, and one rate for each of nine writing systems. The rates are fitted to real bills for the model's tokenizer family, and the fingerprint dataset says which family that is, which is why it is the base of this tool. Every request then adds the fixed overhead the route bills on top of any prompt, measured per model: 8 tokens at the median, and 1,242 on Grok 4.7.
Output tokensThe answer length you set, or the profile's typical figure, at the model's listed completion price.
Thinking tokensThe think bench's median and 90th percentile for the model, the effort and the task type, at the model's reasoning price where the listing prices reasoning separately and the completion price otherwise.
PricesThe OpenRouter listing read on 8 October 2026. A listing that charges more once a prompt passes a set size (41 of them do, Claude Haiku 5.5 five times as much above 100,000 prompt tokens) is priced at that tier for the whole request. A listing priced by time of day is charged its weekly average for requests spread evenly across the week, with the range in the note. A price that moves between snapshots appears in the change history.
FitThe prompt, the answer and the thinking at its 90th percentile have to fit the model's context window, and the answer and thinking its listed answer limit. A model the workload does not fit is listed last, marked does not fit.
RouteWhere a re-measurement found two providers billing the same model differently, the note column says so; the record page carries the strings and the counts.

How far the estimate goes

On 8 October 2026 we billed 18 documents on 12 models, one from each major tokenizer family, and compared each bill with the estimate for the same text. These documents were kept apart from the 24 the rates are fitted to. The input count landed 5.8% from the bill at the median, within 5% on 42% of readings and within 10% on 80%; nine readings in ten were within 16%. Rates read off the 50 fingerprint strings alone, on the same bills, missed by 24.9% at the median.

DocumentMedian missWithin 10%Largest miss
English news article7.3%100%9.8%
English reference page3.2%100%4.9%
English technical README4.9%83%13.9%
English novel3%92%14.4%
TypeScript source1.5%92%16%
Python source7.6%75%15.1%
YAML configuration9.1%58%27.7%
JSON data6.2%100%9.5%
JSON data (minified)5.9%100%9.9%
CSV table11.4%33%20.9%
HTML page markup5.4%92%15.2%
Chinese encyclopedia text3.3%100%7%
Japanese encyclopedia text8.1%75%13.1%
Korean encyclopedia text4.6%92%10.7%
Russian encyclopedia text5.9%100%9.6%
Hindi encyclopedia text4.1%92%25.5%
Spanish encyclopedia text7.5%58%34.4%
German encyclopedia text25.7%0%44%

The prose rate is fitted to English, so other languages written in the Latin alphabet are priced as English, and they bill above the estimate: Spanish text by 7.5% at the median and up to 34.4%, and German text by 25.7% at the median and up to 44%. A pasted sample is sorted line by line, which is right for the mix and approximate at the edges: a code comment counts as prose. The bench tasks are short and self-contained; an agent loop over a long context will think more than the median here, which is one more reason to plan on the 90th percentile and cap it. Prompt caching discounts, batch pricing and volume contracts are left to the buyer's own terms; the table uses the listed price, and a listing priced by time of day at its weekly average. Every figure is dated, and the same table is available as JSON on the data API.