● LAYER 04 OF 04 · OUTPUT
AI model prices: 361 models in dollars per million tokens
Gambit's model registry tracks a blended price per million tokens for 361 models across 50 labs, from Ling-2.6-flash at $0.01/M to o1-pro at $262.50/M, a 26,250× spread, with cost-per-capability analysis on 124 of them.
AS OF 10 AUGUST 2026 · SOURCE: GAMBIT · CURRENT VALUES IN THE TERMINAL ↗
The field
Every lab, counted and priced
| Lab | Models | Median $/M | Cheapest | Priciest |
|---|---|---|---|---|
| OpenAI | 71 | $2.38 | gpt-oss-20b · $0.03 | o1-pro · $262.50 |
| Qwen | 49 | $0.35 | Qwen3.7 Flash · $0.06 | Qwen3.6 Max Preview · $2.31 |
| 37 | $0.56 | Gemma 3 4B · $0.05 | Gemini 3.1 Pro Preview Custom Tools · $4.50 | |
| Anthropic | 26 | $8.00 | Claude 3 Haiku · $0.50 | Claude Opus 4.7 (Fast) · $60.00 |
| Others | 22 | $0.36 | mistral-nemo-instruct-2407 · $0.02 | seed-2-0-code · $1.12 |
| Mistral | 19 | $0.26 | Mistral Nemo · $0.02 | Mistral Medium 3.5 · $3.00 |
| Z.ai | 12 | $0.75 | GLM-4.7-Flash · $0.14 | GLM 5V Turbo · $1.90 |
| DeepSeek | 12 | $0.28 | DeepSeek-V4-Flash · $0.11 | DeepSeek-V4-Pro · $0.54 |
| Meta | 9 | $0.15 | Llama 3.2 3B Instruct · $0.02 | Muse Spark 1.1 · $2.00 |
| MiniMax | 9 | $0.42 | MiniMax M3 (batch) · $0.26 | MiniMax M1 · $0.85 |
| Moonshot | 8 | $1.19 | Kimi K2.5 · $0.79 | Kimi K3 · $5.67 |
| Perplexity | 5 | $3.50 | Sonar · $1.00 | Sonar Pro · $6.00 |
| ByteDance | 5 | $0.17 | UI-TARS 7B · $0.12 | Seed 1.6 · $0.69 |
| Amazon | 5 | $0.85 | Nova Micro 1.0 · $0.06 | Nova Premier 1.0 · $5.00 |
| xAI | 5 | $1.56 | Grok Build 0.1 · $1.25 | Grok 4.5 · $3.00 |
| Aion Labs | 4 | $1.00 | Aion-3.0-Mini · $0.88 | Aion-3.0 · $3.75 |
| TheDrummer | 4 | $0.38 | Rocinante 12B · $0.31 | Skyfall 36B V2 · $0.61 |
| Cohere | 4 | $2.32 | Command R7B (12-2024) · $0.07 | Command A · $4.38 |
| Nous Research | 4 | $0.60 | Hermes 3 70B Instruct · $0.17 | Hermes 4 405B · $1.50 |
| Inclusionai | 3 | $0.21 | Ling-2.6-flash · $0.01 | Ling-2.6-1T · $0.21 |
| Nvidia | 3 | $0.16 | Nemotron 3 Nano 30B A3B · $0.09 | Nemotron 3 Ultra · $0.93 |
| Sao10K | 3 | $0.68 | Llama 3 8B Lunaris · $0.04 | Llama 3.1 Euryale 70B v2.2 · $0.85 |
| Tencent | 3 | $0.23 | Hy3-preview · $0.10 | Hunyuan-A13B-Instruct · $0.25 |
| Kwaipilot | 3 | $0.53 | KAT-Coder-Air V2.5 · $0.26 | KAT-Coder-Pro V2.5 · $1.29 |
| Xiaomi MiMo | 2 | $0.29 | MiMo-V2.5 · $0.14 | MiMo-V2.5-Pro · $0.43 |
| Arcee | 2 | $0.62 | Trinity Large Thinking · $0.38 | Virtuoso Large · $0.86 |
| IBM Granite | 2 | $0.05 | granite-4.0-h-micro · $0.04 | granite-4.1-8b · $0.06 |
| StepFun | 2 | $0.29 | Step-3.5-Flash · $0.15 | Step 3.7 Flash · $0.44 |
| Reka | 2 | $0.11 | Reka Edge · $0.10 | Reka Flash 3 · $0.12 |
| Poolside | 2 | $0.09 | Laguna-XS-2.1 · $0.07 | Laguna-S-2.1 · $0.11 |
| Nex-Agi | 2 | $0.24 | Nex-N2-Mini · $0.04 | Nex-N2-Pro · $0.44 |
| Microsoft | 2 | $0.28 | phi-4 · $0.09 | WizardLM-2 8x22B · $0.48 |
| Relace | 2 | $1.23 | Relace Apply 3 · $0.95 | Relace Search · $1.50 |
| Morph | 2 | $1.02 | Morph V3 Fast · $0.90 | Morph V3 Large · $1.15 |
| Inception | 1 | $0.38 | Mercury 2 · $0.38 | Mercury 2 · $0.38 |
| Gryphe | 1 | $0.06 | MythoMax 13B · $0.06 | MythoMax 13B · $0.06 |
| Deepcogito | 1 | $1.25 | Cogito v2.1 671B · $1.25 | Cogito v2.1 671B · $1.25 |
| Cognitive Computations | 1 | $0.38 | Uncensored · $0.38 | Uncensored · $0.38 |
| Baidu | 1 | $0.63 | ERNIE 4.5 VL 424B A47B · $0.63 | ERNIE 4.5 VL 424B A47B · $0.63 |
| Anthracite | 1 | $3.50 | Magnum v4 72B · $3.50 | Magnum v4 72B · $3.50 |
| AI21 | 1 | $3.50 | Jamba Large 1.7 · $3.50 | Jamba Large 1.7 · $3.50 |
| Meituan | 1 | $0.53 | LongCat 2.0 · $0.53 | LongCat 2.0 · $0.53 |
| Perceptron | 1 | $0.49 | Perceptron Mk1 · $0.49 | Perceptron Mk1 · $0.49 |
| Mancer | 1 | $0.56 | Weaver (alpha) · $0.56 | Weaver (alpha) · $0.56 |
| Sakana | 1 | $11.25 | Fugu Ultra · $11.25 | Fugu Ultra · $11.25 |
| Thinking Machines | 1 | $1.76 | Inkling · $1.76 | Inkling · $1.76 |
| Upstage | 1 | $0.26 | Solar Pro 3 · $0.26 | Solar Pro 3 · $0.26 |
| Unsloth | 1 | $0.65 | gemma-2-27b-it · $0.65 | gemma-2-27b-it · $0.65 |
| Undi95 | 1 | $0.50 | ReMM SLERP 13B · $0.50 | ReMM SLERP 13B · $0.50 |
| Writer | 1 | $1.95 | Palmyra X5 · $1.95 | Palmyra X5 · $1.95 |
361 models across 50 labs are priced on the DATA desk, with cost-per-capability on 124 of them.
Compare all 361 models in the terminal ↗The extremes
The 10 cheapest and the 10 priciest
| Model | Lab | $/M tokens |
|---|---|---|
| Ling-2.6-flash | Inclusionai | $0.01 |
| Mistral Nemo | Mistral | $0.02 |
| Llama 3.2 3B Instruct | Meta | $0.02 |
| gpt-oss-20b | OpenAI | $0.03 |
| Llama 3.1 8B Instruct | Meta | $0.03 |
| Llama 3 8B Lunaris | Sao10K | $0.04 |
| granite-4.0-h-micro | IBM Granite | $0.04 |
| Nex-N2-Mini | Nex-Agi | $0.04 |
| Gemma 3 4B | $0.05 | |
| Qwen3.7 Flash | Qwen | $0.06 |
| Model | Lab | $/M tokens |
|---|---|---|
| o1-pro | OpenAI | $262.50 |
| Claude Opus 4.7 (Fast) | Anthropic | $60.00 |
| GPT-5.2 Pro | OpenAI | $57.75 |
| GPT-5 Pro | OpenAI | $41.25 |
| GPT-4 | OpenAI | $37.50 |
| o3 Pro | OpenAI | $35.00 |
| GPT-5.5 Pro | OpenAI | $33.75 |
| GPT-5.4 Pro | OpenAI | $33.75 |
| Claude Opus 4 | Anthropic | $30.00 |
| Claude Opus 4.1 | Anthropic | $30.00 |
Price does not track recency. Providers rarely mark a legacy endpoint down when its successor ships, so the top of this table mixes current flagship tiers with superseded models that were simply never repriced. Today GPT-4 lists at $37.50 while GPT-5.5 Pro lists at $33.75: the older model is the more expensive one. Read the top of this board as what you overpay by not migrating, not as what frontier capability costs.
These extremes are drawn from the 339 models whose lab the registry resolves. 22 further rows are merchant-specific slugs, usually a duplicate listing of a model already named above, and are counted in the lab table but kept out of the rankings. The full sortable board, with cost-per-capability scores, is on the DATA desk.
Questions
How many AI models does Gambit track prices for?
361 models across 50 labs, each with a blended price per million tokens; a cost-per-capability analysis covers 124 of them. The registry spans frontier APIs (OpenAI, Anthropic, Google) and open-weight labs (Qwen, DeepSeek, Meta, Mistral). As of 10 August 2026.
What is the cheapest LLM API right now?
At this snapshot the lowest blended prices tracked are Ling-2.6-flash (Inclusionai) at $0.01/M tokens, Mistral Nemo (Mistral) at $0.02/M tokens, Llama 3.2 3B Instruct (Meta) at $0.02/M tokens. The most expensive is o1-pro (OpenAI) at $262.50/M, a 26,250× spread across the registry.
What does "blended $/M tokens" mean?
A single per-million-token figure combining a model’s input and output token prices, as tracked by the DATA desk, so models with very different input/output ratios can sit in one sortable column.
Why do model prices matter for the AI trade?
Token prices are the revenue side of the GPU economy: what a chip earns per hour must ultimately be paid for by what its tokens sell for. Read this page against GPU rental prices and hyperscaler capex to see all three layers of the stack priced.
The live desk
What this page leaves out
This page publishes the per-lab summary and the ten models at each end of the price range. The desk holds the whole registry, the cost-per-capability score where evaluation data exists, and the individual merchant listings and source links behind every price.
- 361models priced, against 20 published here
- 124with a cost-per-capability score
- 1,048individual merchant listings behind the prices
- 1,770model spec records for context
Back to the top of the chain
Hyperscaler capex
Compare capital growth, operational capacity, compute prices and model pricing together before drawing a conclusion about the cycle. The chain starts again at what is being committed: Is investment accelerating?