Runyard / LLM API Index · GPU Index →

LLM API Pricing Index

Input and output price per million tokens for 440 models from 59 vendors, in one schema. 418 are paid, 22 have a free tier. Blended price assumes 3 input tokens per output token, the ratio most chat and agent work lands near.

Last collected 2026-09-12 · 7 snapshots

Models priced
44059 vendors
Median paid model
$0.850/1M blended
Frontier tier
$20.00/1M blended
Cheapest paid
$0.022Mistral Nemo
Flagships, blended $/1M tokens
3 input : 1 output · as listed on OpenRouter
Claude Fable 5.1$20.00GPT-6 Astra$20.00Grok 4.6$3.00Kimi K3$5.31DeepSeek V4.1 Flash$0.26GLM 5.3 Flash$0.24Gemini 3.8 Flash$1.50

The spread is the story: the frontier tier costs 13× the cheapest flagship-class model on this chart, for work that often does not need it.

Every model

ContextCache read7D
Mistral NemoMistral · mistralai/mistral-nemo131K$0.019$0.030$0.022
Ling 3.0 FlashinclusionAI · inclusionai/ling-3.0-flash262K$0.021$0.063$0.032$0.0042
DeepSeek V4 Flash Latest~Deepseek · ~deepseek/deepseek-v4-flash-latest1.3M$0.030$0.070$0.040$0.0030
Granite 4.0 MicroIbm Granite · ibm-granite/granite-4.0-h-micro131K$0.017$0.112$0.041
Llama 3 8B LunarisSao10k · sao10k/l3-lunaris-8b8K$0.040$0.050$0.043
DeepSeek V4 Flash 0731DeepSeek · deepseek/deepseek-v4-flash-07311.3M$0.040$0.080$0.050$0.0080
Qwen3.7 FlashAlibaba Qwen · qwen/qwen3.7-flash1M$0.030$0.130$0.055$0.0060
gpt-oss-20bOpenAI · openai/gpt-oss-20b131K$0.030$0.130$0.055$0.030
Mistral Small 3Mistral · mistralai/mistral-small-24b-instruct-250133K$0.050$0.080$0.058
Llama 3.1 8B InstructMeta · meta-llama/llama-3.1-8b-instruct131K$0.050$0.080$0.058$0.025
Schematron V2 TurboInference Net · inference-net/schematron-v2-turbo128K$0.030$0.150$0.060
MythoMax 13BGryphe · gryphe/mythomax-l2-13b8K$0.060$0.060$0.060
Nova Micro 1.0Amazon · amazon/nova-micro-v1128K$0.035$0.140$0.061
Gemma 3 4BGoogle · google/gemma-3-4b-it131K$0.050$0.100$0.063
Command R7B (12-2024)Cohere · cohere/command-r7b-12-2024128K$0.037$0.150$0.066
Mercury 2.5Inception · inception/mercury-2.5260K$0.040$0.150$0.068$0.0040
GPT-5 Nano (batch)OpenAI · openai/gpt-5-nano:batch400K$0.025$0.200$0.069$0.0025
gpt-oss-120bOpenAI · openai/gpt-oss-120b131K$0.037$0.170$0.070
Llama 3.2 1B InstructMeta · meta-llama/llama-3.2-1b-instruct60K$0.027$0.201$0.070
Laguna XS 2.1Poolside · poolside/laguna-xs-2.1262K$0.060$0.120$0.075$0.030
Ministral 3 8B 2512 (batch)Mistral · mistralai/ministral-8b-2512:batch262K$0.075$0.075$0.075$0.0075
Gemma 3 12BGoogle · google/gemma-3-12b-it131K$0.050$0.150$0.075
Hy-MT2-1.8BTencent · tencent/hy-mt2-1.8b8K$0.044$0.177$0.077
DeepSeek V4 Flash 0423DeepSeek · deepseek/deepseek-v4-flash1.0M$0.066$0.131$0.082$0.013
Gemma 4 26B A4B Google · google/gemma-4-26b-a4b-it262K$0.042$0.220$0.086
Nemotron 3 Nano 30B A3BNVIDIA · nvidia/nemotron-3-nano-30b-a3b262K$0.050$0.200$0.087$0.030
gpt-oss-20b (batch)OpenAI · openai/gpt-oss-20b:batch131K$0.050$0.200$0.087
Gemini 2.5 Flash Lite (batch)Google · google/gemini-2.5-flash-lite:batch1.0M$0.050$0.200$0.087$0.010
GPT-4.1 Nano (batch)OpenAI · openai/gpt-4.1-nano:batch1.0M$0.050$0.200$0.087$0.013
Phi 4Microsoft · microsoft/phi-416K$0.070$0.140$0.087
Ling 3.0 Flash VLinclusionAI · inclusionai/ling-3.0-flash-vl131K$0.060$0.180$0.090$0.012
Ling 3.0 Flash FininclusionAI · inclusionai/ling-3.0-flash-fin262K$0.060$0.180$0.090$0.012
Schematron V2 SmallInference Net · inference-net/schematron-v2-small128K$0.050$0.230$0.095
Reka EdgeRekaai · rekaai/reka-edge16K$0.100$0.100$0.100
Ministral 3 3B 2512Mistral · mistralai/ministral-3b-2512131K$0.100$0.100$0.100$0.010
Nova Lite 1.0Amazon · amazon/nova-lite-v1300K$0.060$0.240$0.105
Mistral Small 3.2 24BMistral · mistralai/mistral-small-3.2-24b-instruct256K$0.075$0.200$0.106
Granite 4.2 8BIbm Granite · ibm-granite/granite-4.2-8b131K$0.060$0.250$0.107$0.015
Nemotron 3.5 LightningNVIDIA · nvidia/nemotron-3.5-lightning262K$0.080$0.200$0.110$0.040
Laguna S 2.1Poolside · poolside/laguna-s-2.11.0M$0.090$0.180$0.113$0.0090
Qwen3.5-9BAlibaba Qwen · qwen/qwen3.5-9b262K$0.100$0.150$0.113
Qwen3.5-FlashAlibaba Qwen · qwen/qwen3.5-flash-02-231M$0.065$0.260$0.114
GLM Flash Latest~Z Ai · ~z-ai/glm-flash-latest1.3M$0.075$0.250$0.119$0.015
GLM 5.3 Flash (batch)Z.ai (Zhipu) · z-ai/glm-5.3-flash:batch1.0M$0.075$0.250$0.119$0.015
Llama 3.2 3B InstructMeta · meta-llama/llama-3.2-3b-instruct131K$0.050$0.330$0.120
Qwen3 Coder 30B A3B InstructAlibaba Qwen · qwen/qwen3-coder-30b-a3b-instruct262K$0.070$0.280$0.122
Muse Spark 1.3 ContributorMeta · meta/muse-spark-1.3-contributor1.0M$0.100$0.200$0.125$0.0020
Muse Spark 1.2 ContributorMeta · meta/muse-spark-1.2-contributor1.0M$0.100$0.200$0.125$0.0020
UI-TARS 7B Bytedance · bytedance/ui-tars-1.5-7b128K$0.100$0.200$0.125$0.100
Reka Flash 3Rekaai · rekaai/reka-flash-366K$0.100$0.200$0.125
Qwen2.5 7B InstructAlibaba Qwen · qwen/qwen-2.5-7b-instruct33K$0.100$0.200$0.125
Hy-MT2-30B-A3BTencent · tencent/hy-mt2-30b-a3b8K$0.074$0.295$0.129
Hy-MT2-7BTencent · tencent/hy-mt2-7b8K$0.074$0.295$0.129
Qwen3 32BAlibaba Qwen · qwen/qwen3-32b131K$0.080$0.280$0.130
Mistral Small 4 (batch)Mistral · mistralai/mistral-small-2603:batch262K$0.075$0.300$0.131$0.0075
Seed 1.6 FlashBytedance Seed · bytedance-seed/seed-1.6-flash262K$0.075$0.300$0.131
gpt-oss-safeguard-20bOpenAI · openai/gpt-oss-safeguard-20b131K$0.075$0.300$0.131$0.037
GPT-4o-mini (batch)OpenAI · openai/gpt-4o-mini:batch128K$0.075$0.300$0.131$0.037
GPT-5 NanoOpenAI · openai/gpt-5-nano400K$0.050$0.400$0.138$0.0050
Qwen3 30B A3B Instruct 2507Alibaba Qwen · qwen/qwen3-30b-a3b-instruct-2507262K$0.090$0.300$0.142
Hy3Tencent · tencent/hy3262K$0.083$0.330$0.144$0.021
GLM 4.7 FlashZ.ai (Zhipu) · z-ai/glm-4.7-flash200K$0.060$0.400$0.145
Step 3.5 FlashStepfun · stepfun/step-3.5-flash262K$0.100$0.300$0.150
Ministral 3 8B 2512Mistral · mistralai/ministral-8b-2512262K$0.150$0.150$0.150$0.015
Voxtral Small 24B 2507Mistral · mistralai/voxtral-small-24b-250733K$0.100$0.300$0.150$0.010
Llama 4 ScoutMeta · meta-llama/llama-4-scout1.3M$0.100$0.300$0.150
Gemma 4 31BGoogle · google/gemma-4-31b-it262K$0.090$0.340$0.152$0.050
Qwen3 235B A22B Instruct 2507Alibaba Qwen · qwen/qwen3-235b-a22b-2507262K$0.087$0.350$0.153$0.018
Llama 3.3 70B InstructMeta · meta-llama/llama-3.3-70b-instruct131K$0.100$0.320$0.155
Solar Pro 4Upstage · upstage/solar-pro4524K$0.090$0.360$0.158$0.018
Nemotron 3 SuperNVIDIA · nvidia/nemotron-3-super-120b-a12b262K$0.085$0.400$0.164
DeepSeek V4 Flash Vision Exp (batch)DeepSeek · deepseek/deepseek-v4-flash-vision-exp:batch1.0M$0.110$0.330$0.165$0.0035
DeepSeek V4 Flash 0731 (batch)DeepSeek · deepseek/deepseek-v4-flash-0731:batch1.0M$0.110$0.330$0.165$0.0035
Gemma 3 27BGoogle · google/gemma-3-27b-it131K$0.080$0.450$0.172$0.040
MiMo-V2.5Xiaomi · xiaomi/mimo-v2.51.1M$0.140$0.280$0.175$0.0028
Seed-2.0-MiniBytedance Seed · bytedance-seed/seed-2.0-mini262K$0.100$0.400$0.175
Gemini 2.5 Flash LiteGoogle · google/gemini-2.5-flash-lite1.0M$0.100$0.400$0.175$0.010
GPT-4.1 NanoOpenAI · openai/gpt-4.1-nano1.0M$0.100$0.400$0.175$0.025
Llama Guard 4 12BMeta · meta-llama/llama-guard-4-12b164K$0.180$0.180$0.180
Qwen3 VL 32B InstructAlibaba Qwen · qwen/qwen3-vl-32b-instruct131K$0.104$0.416$0.182

80 of 418 models (free tiers hidden). Blended = 3 input tokens per output token. Prices as listed on OpenRouter on the date above.

Vendors

VendorModelsCheapest paid, blended
OpenAI93$0.055
Alibaba Qwen53$0.055
Google43$0.063
Anthropic27$0.500
Mistral25$0.022
DeepSeek18$0.050
Z.ai (Zhipu)17$0.119
NVIDIA10$0.087
MiniMax9$0.425
Moonshot AI8$0.900
Meta8$0.058
Meta7$0.125
Tencent7$0.077
xAI7$1.25
inclusionAI6$0.032
Bytedance Seed6$0.131
Thinking Machines6$0.637
~Openai5$0.450
Cohere5$0.066
Amazon5$0.061
Perplexity5$1.00
Sakana AI4$1.71
Poolside4$0.075
Aion Labs4$0.875
~Anthropic4$2.00
Thedrummer3$0.350
Nous Research3$0.700
Sao10k3$0.043
Inference Net2$0.060
Inception2$0.068
Nex Agi2
Ibm Granite2$0.041
~Z Ai2$0.119
Upstage2$0.158
Kwaipilot2$0.525
Stepfun2$0.150
~Google2$1.50
Xiaomi2$0.175
Rekaai2$0.100
Relace2$0.950
Morph2$0.900
Microsoft2$0.087
Dots Studio1
Liquid AI1
~Deepseek1$0.040
Meituan1$0.525
~X Ai1$3.00
Perceptron1$0.487
~Moonshotai1$5.11
Arcee Ai1$0.388
Openrouter1
Writer1$1.95
Bytedance1$0.125
Cognitivecomputations1$0.375
Baidu1$0.627
Anthracite Org1$3.13
Mancer1$0.487
Undi951$0.425
Gryphe1$0.060

How this is built

Prices are read from OpenRouter's public model list, which routes to each vendor and publishes the per-token price it charges — for nearly every model, the vendor's own list price passed through. It is the one place all vendors appear in one schema, which is why it is the source. Prices are shown as listed there; a vendor's direct price can differ, especially with volume discounts or provisioned throughput, and OpenRouter's own listing is the reference for what is shown.

Every collector run is stored as a snapshot and committed to the repository, so price changes are diffs anyone can inspect. Nothing is backfilled. The 7-day column compares a model against itself a week earlier and is blank until that run exists.

Blended weights input at 3 and output at 1. Change the ratio and the ranking changes — an output-heavy workload makes the frontier tier relatively more expensive, since output tokens cost five times input on most frontier models.

What the same tokens cost on hardware you rent or own.
Local cost per million tokens →