GAMBIT_TERMINAL OPEN THE TERMINAL ↗

LAYER 04 OF 04 · OUTPUT

AI model prices: 361 models in dollars per million tokens

Gambit's model registry tracks a blended price per million tokens for 361 models across 50 labs, from Ling-2.6-flash at $0.01/M to o1-pro at $262.50/M, a 26,250× spread, with cost-per-capability analysis on 124 of them.

AS OF 10 AUGUST 2026 · SOURCE: GAMBIT · CURRENT VALUES IN THE TERMINAL ↗

The field

Every lab, counted and priced

Models tracked per lab with median, cheapest and priciest blended $/M tokens. Labs with the most tracked models first.
LabModelsMedian $/MCheapestPriciest
OpenAI71$2.38gpt-oss-20b · $0.03o1-pro · $262.50
Qwen49$0.35Qwen3.7 Flash · $0.06Qwen3.6 Max Preview · $2.31
Google37$0.56Gemma 3 4B · $0.05Gemini 3.1 Pro Preview Custom Tools · $4.50
Anthropic26$8.00Claude 3 Haiku · $0.50Claude Opus 4.7 (Fast) · $60.00
Others22$0.36mistral-nemo-instruct-2407 · $0.02seed-2-0-code · $1.12
Mistral19$0.26Mistral Nemo · $0.02Mistral Medium 3.5 · $3.00
Z.ai12$0.75GLM-4.7-Flash · $0.14GLM 5V Turbo · $1.90
DeepSeek12$0.28DeepSeek-V4-Flash · $0.11DeepSeek-V4-Pro · $0.54
Meta9$0.15Llama 3.2 3B Instruct · $0.02Muse Spark 1.1 · $2.00
MiniMax9$0.42MiniMax M3 (batch) · $0.26MiniMax M1 · $0.85
Moonshot8$1.19Kimi K2.5 · $0.79Kimi K3 · $5.67
Perplexity5$3.50Sonar · $1.00Sonar Pro · $6.00
ByteDance5$0.17UI-TARS 7B · $0.12Seed 1.6 · $0.69
Amazon5$0.85Nova Micro 1.0 · $0.06Nova Premier 1.0 · $5.00
xAI5$1.56Grok Build 0.1 · $1.25Grok 4.5 · $3.00
Aion Labs4$1.00Aion-3.0-Mini · $0.88Aion-3.0 · $3.75
TheDrummer4$0.38Rocinante 12B · $0.31Skyfall 36B V2 · $0.61
Cohere4$2.32Command R7B (12-2024) · $0.07Command A · $4.38
Nous Research4$0.60Hermes 3 70B Instruct · $0.17Hermes 4 405B · $1.50
Inclusionai3$0.21Ling-2.6-flash · $0.01Ling-2.6-1T · $0.21
Nvidia3$0.16Nemotron 3 Nano 30B A3B · $0.09Nemotron 3 Ultra · $0.93
Sao10K3$0.68Llama 3 8B Lunaris · $0.04Llama 3.1 Euryale 70B v2.2 · $0.85
Tencent3$0.23Hy3-preview · $0.10Hunyuan-A13B-Instruct · $0.25
Kwaipilot3$0.53KAT-Coder-Air V2.5 · $0.26KAT-Coder-Pro V2.5 · $1.29
Xiaomi MiMo2$0.29MiMo-V2.5 · $0.14MiMo-V2.5-Pro · $0.43
Arcee2$0.62Trinity Large Thinking · $0.38Virtuoso Large · $0.86
IBM Granite2$0.05granite-4.0-h-micro · $0.04granite-4.1-8b · $0.06
StepFun2$0.29Step-3.5-Flash · $0.15Step 3.7 Flash · $0.44
Reka2$0.11Reka Edge · $0.10Reka Flash 3 · $0.12
Poolside2$0.09Laguna-XS-2.1 · $0.07Laguna-S-2.1 · $0.11
Nex-Agi2$0.24Nex-N2-Mini · $0.04Nex-N2-Pro · $0.44
Microsoft2$0.28phi-4 · $0.09WizardLM-2 8x22B · $0.48
Relace2$1.23Relace Apply 3 · $0.95Relace Search · $1.50
Morph2$1.02Morph V3 Fast · $0.90Morph V3 Large · $1.15
Inception1$0.38Mercury 2 · $0.38Mercury 2 · $0.38
Gryphe1$0.06MythoMax 13B · $0.06MythoMax 13B · $0.06
Deepcogito1$1.25Cogito v2.1 671B · $1.25Cogito v2.1 671B · $1.25
Cognitive Computations1$0.38Uncensored · $0.38Uncensored · $0.38
Baidu1$0.63ERNIE 4.5 VL 424B A47B · $0.63ERNIE 4.5 VL 424B A47B · $0.63
Anthracite1$3.50Magnum v4 72B · $3.50Magnum v4 72B · $3.50
AI211$3.50Jamba Large 1.7 · $3.50Jamba Large 1.7 · $3.50
Meituan1$0.53LongCat 2.0 · $0.53LongCat 2.0 · $0.53
Perceptron1$0.49Perceptron Mk1 · $0.49Perceptron Mk1 · $0.49
Mancer1$0.56Weaver (alpha) · $0.56Weaver (alpha) · $0.56
Sakana1$11.25Fugu Ultra · $11.25Fugu Ultra · $11.25
Thinking Machines1$1.76Inkling · $1.76Inkling · $1.76
Upstage1$0.26Solar Pro 3 · $0.26Solar Pro 3 · $0.26
Unsloth1$0.65gemma-2-27b-it · $0.65gemma-2-27b-it · $0.65
Undi951$0.50ReMM SLERP 13B · $0.50ReMM SLERP 13B · $0.50
Writer1$1.95Palmyra X5 · $1.95Palmyra X5 · $1.95

361 models across 50 labs are priced on the DATA desk, with cost-per-capability on 124 of them.

Compare all 361 models in the terminal ↗

The extremes

The 10 cheapest and the 10 priciest

The 10 lowest blended prices tracked, $/M tokens.
ModelLab$/M tokens
Ling-2.6-flashInclusionai$0.01
Mistral NemoMistral$0.02
Llama 3.2 3B InstructMeta$0.02
gpt-oss-20bOpenAI$0.03
Llama 3.1 8B InstructMeta$0.03
Llama 3 8B LunarisSao10K$0.04
granite-4.0-h-microIBM Granite$0.04
Nex-N2-MiniNex-Agi$0.04
Gemma 3 4BGoogle$0.05
Qwen3.7 FlashQwen$0.06
The 10 highest blended prices tracked, $/M tokens.
ModelLab$/M tokens
o1-proOpenAI$262.50
Claude Opus 4.7 (Fast)Anthropic$60.00
GPT-5.2 ProOpenAI$57.75
GPT-5 ProOpenAI$41.25
GPT-4OpenAI$37.50
o3 ProOpenAI$35.00
GPT-5.5 ProOpenAI$33.75
GPT-5.4 ProOpenAI$33.75
Claude Opus 4Anthropic$30.00
Claude Opus 4.1Anthropic$30.00

Price does not track recency. Providers rarely mark a legacy endpoint down when its successor ships, so the top of this table mixes current flagship tiers with superseded models that were simply never repriced. Today GPT-4 lists at $37.50 while GPT-5.5 Pro lists at $33.75: the older model is the more expensive one. Read the top of this board as what you overpay by not migrating, not as what frontier capability costs.

These extremes are drawn from the 339 models whose lab the registry resolves. 22 further rows are merchant-specific slugs, usually a duplicate listing of a model already named above, and are counted in the lab table but kept out of the rankings. The full sortable board, with cost-per-capability scores, is on the DATA desk.

Questions

How many AI models does Gambit track prices for?

361 models across 50 labs, each with a blended price per million tokens; a cost-per-capability analysis covers 124 of them. The registry spans frontier APIs (OpenAI, Anthropic, Google) and open-weight labs (Qwen, DeepSeek, Meta, Mistral). As of 10 August 2026.

What is the cheapest LLM API right now?

At this snapshot the lowest blended prices tracked are Ling-2.6-flash (Inclusionai) at $0.01/M tokens, Mistral Nemo (Mistral) at $0.02/M tokens, Llama 3.2 3B Instruct (Meta) at $0.02/M tokens. The most expensive is o1-pro (OpenAI) at $262.50/M, a 26,250× spread across the registry.

What does "blended $/M tokens" mean?

A single per-million-token figure combining a model’s input and output token prices, as tracked by the DATA desk, so models with very different input/output ratios can sit in one sortable column.

Why do model prices matter for the AI trade?

Token prices are the revenue side of the GPU economy: what a chip earns per hour must ultimately be paid for by what its tokens sell for. Read this page against GPU rental prices and hyperscaler capex to see all three layers of the stack priced.

The live desk

What this page leaves out

This page publishes the per-lab summary and the ten models at each end of the price range. The desk holds the whole registry, the cost-per-capability score where evaluation data exists, and the individual merchant listings and source links behind every price.

Back to the top of the chain

Hyperscaler capex

Compare capital growth, operational capacity, compute prices and model pricing together before drawing a conclusion about the cycle. The chain starts again at what is being committed: Is investment accelerating?