← /blog
· ACE Engineering#llm #pricing #model-catalog #market #cost #benchmarks #managed-api-stack

The 2026 model market, in one table: every model, provider, and price worth comparing

A versioned, sourced catalog of the models on the market today — per-token in/out prices, cache-read rates, context, modalities, capability evidence, and the Artificial Analysis Intelligence Index — grouped by family so you can compare price against performance at a glance.

The 2026 model market, in one table

Catalog version 2026.07.2 · compiled 2026-07-24. This is a snapshot of the models you can buy today, the providers that serve them, and what they cost — assembled so you can compare price against performance without cross-referencing a dozen pricing pages and leaderboards.

Two models can sit in the same tier of a public coding benchmark and differ by ~34× on the output token price. GPT-5.6 Sol lists at $5.00 / $30.00 per million tokens (in / out); DeepSeek V4 Pro lists at $0.435 / $0.87. The gap between "the name you trust" and "the price you pay" is the entire reason to keep a table like this.

How to read the tables

  • Price is per 1M tokens, input / output. Cache read is the discounted rate for prompt-cache hits where a provider publishes one.
  • Capability scores are 0..1 per task category, and each carries a basis you should weigh:
    • published — a cited benchmark measurement with a URL and date (e.g. GPQA Diamond 94.1%).
    • op-rank — the model's placement in an operator's July-2026 shortlist for that category: directional, not a number it earned on a board we hold.
    • assumed — a fallback estimate.
    • Most capability placements below are operator-supplied and not independently verified — several of these models postdate our sourcing window. Prices are cited; capability rankings largely are not yet. Cells are left blank rather than invented.
  • AA Idx is the Artificial Analysis Intelligence Index (v4.1 composite of 9 evaluations, as of 2026-07-23) — a single model-level "general capability" number for cross-family comparison, not a per-task score.

Anthropic — Claude

Source: claude.com/pricing and platform.claude.com pricing — first-party prices confirmed 2026-07-24.

Model?In $/1M?Out $/1M?Cache read?Context?Capability (basis)?AA Idx?
claude-fable-5 10.00 50.00 1.00 0.95code/reasoning/qa/visionop-rank0.90chat 60
claude-mythos-5 10.00 50.00 1.00 1M
claude-opus-4-8 5.00 25.00 0.50 0.95code0.95long_contextop-rank0.90chat 56
claude-opus-4-7 5.00 25.00 0.50 1M
claude-opus-4-6 5.00 25.00 0.50 1M
claude-opus-4-5 5.00 25.00 0.50 200k 0.895qapublished, MMLU-Pro 89.5%
claude-opus-4-1 15.00 75.00 1.50 200k (deprecated)
claude-sonnet-5 2.00 10.00 0.20 0.88vision/codeop-rank0.88chat 53
claude-sonnet-4-6 3.00 15.00 0.30 1M
claude-sonnet-4-5 3.00 15.00 0.30 200k
claude-haiku-4-5 1.00 5.00 0.10 200k
claude-haiku-3-5 0.80 4.00 0.08 200k (retired; Bedrock/GCP only)

claude-sonnet-5 is at an introductory rate through 2026-08-31 ($2/$10, $0.20 cache read); it reverts to $3/$15 ($0.30 cache read) on 2026-09-01. All Claude models are text+vision except Haiku 3.5 and Mythos 5 (text).

OpenAI — GPT-5

Source: developers.openai.com pricing, as of 2026-07-23. Also served on Azure OpenAI at parity for gpt-5.6-sol.

Model?In $/1M?Out $/1M?Capability (basis)?AA Idx?Notes?
gpt-5.6-sol 5.00 30.00 0.95math/code/qaop-rank0.941reasoningpublished, GPQA 94.1%0.88vision 59 flagship
gpt-5.6-luna 1.00 6.00 text+vision
gpt-5.6-terra 2.50 15.00 text+vision
gpt-5.5 5.00 30.00
gpt-5.5-pro 30.00 180.00 pro tier
gpt-5.4 2.50 15.00
gpt-5.4-pro 30.00 180.00 pro tier
gpt-5.4-mini 0.75 4.50 text+vision
gpt-5.4-nano 0.20 1.25 text
gpt-5.3-codex 1.75 14.00 code-specialised, text

The pro tiers at $30 / $180 are the most expensive offerings in the catalog — roughly 6× the flagship on the input leg, 6× on output.

Google — Gemini

Source: ai.google.dev pricing, as of 2026-07-23, except the two pro/flash entries marked Vertex (operator table, 2026-07-21, unverified). All are text+vision.

Model?In $/1M?Out $/1M?Context?Capability (basis)?AA Idx?Served via?
gemini-3.1-pro 2.00 12.00 0.95math/long_contextop-rank0.941reasoningpublished, GPQA 94.1%0.88qa 46 Vertex
gemini-3.6-flash 1.50 7.50 0.88long_context/summarization/extractionop-rank 50 Vertex
gemini-3.5-flash 1.50 9.00 Gemini API
gemini-3.5-flash-lite 0.30 2.50 Gemini API
gemini-3.1-flash-lite 0.25 1.50 Gemini API
gemini-2.5-pro 1.25 10.00 ≤200k Gemini API
gemini-2.5-flash 0.30 2.50 1M Gemini API
gemini-2.5-flash-lite 0.10 0.40 Gemini API

gemini-2.5-flash-lite at $0.10 / $0.40 is the cheapest metered offering in the entire catalog.

DeepSeek — open-weight, first-party API

Source: operator table, 2026-07-21 (prices cited, capability placements unverified). All three are open-weight and served on DeepSeek's OpenAI-compatible API; V4 Pro is also hosted on Fireworks and Together, which quoted no public per-token price.

Model?In $/1M?Out $/1M?Capability (basis)?AA Idx?
deepseek-v4-pro 0.435 0.87 0.88code/extraction/summarizationop-rank0.85chat 44
deepseek-v4-base 0.27 0.55 0.88mathop-rank0.80chat
deepseek-r1 0.55 2.19 0.88math/reasoningop-rank0.80chat

At $0.435 / $0.87, deepseek-v4-pro is roughly 1/10th the price of the flagships it shares a benchmark tier with for coding.

Specialists & other open-weight

Source: operator table, 2026-07-21 (unverified) except where a published board is cited.

Model?Provider(s)?In $/1M?Out $/1M?Capability (basis)?AA Idx?
kimi-k3 Moonshot / OpenRouter 3.00 15.00 0.935reasoningpublished, GPQA 93.5%0.85chat 57
glm-5.2 Zhipu / OpenRouter 1.40 4.40 0.85chat/summarizationop-rank 51
llama-3-70b Self-hosted / Fireworks serving-derived serving-derived 0.80chat/extraction/summarizationassumed

llama-3-70b has no single list price: bought as committed capacity (e.g. Azure PTU) its marginal cost is near-zero until a utilization knee; run on your own H100s its per-token price is derived from instance $/hr divided by achieved throughput; rented serverless it is metered like any vendor API. Same weights, three different cost curves — which is why the market is a matrix, not a list.


The price-performance reading

Sort the evidenced models by the cheapest option that carries a real capability score for each kind of work, against a flagship in the same tier:

Workload?Cheapest with capability evidence?Its price (in/out)?Flagship comparison?Flagship price?
Coding deepseek-v4-pro (0.88) 0.435 / 0.87 gpt-5.6-sol · claude-opus-4-8 5 / 30 · 5 / 25
Math deepseek-v4-base (0.88) 0.27 / 0.55 gpt-5.6-sol 5 / 30
Reasoning deepseek-r1 (0.88) 0.55 / 2.19 kimi-k3 (0.935) · gpt-5.6-sol (0.941) 3 / 15 · 5 / 30
Extraction / summarization deepseek-v4-pro (0.88) 0.435 / 0.87 claude-opus-4-8 5 / 25
Chat deepseek-v4-base (0.80) 0.27 / 0.55 claude-fable-5 (0.90) 10 / 50
QA claude-opus-4-5 (0.895) 5 / 25 claude-fable-5 / gpt-5.6-sol (0.95) 10 / 50 · 5 / 30
Translation (no model in catalog carries a translation score)

The blank rows are the honest part: for QA there is no cheap model with published or shortlisted evidence, and for translation there is no capability number at all. A catalog that filled those in would be inventing data.


Provenance and staleness

  • Anthropic first-party prices: vendor-confirmed 2026-07-24.
  • OpenAI / Google prices: vendor docs, as-of 2026-07-23.
  • DeepSeek, Kimi, GLM, GPT-5.6 Sol, Gemini 3.x capability placements and several prices: operator shortlist, as-of 2026-07-21, independently unverified — treat as directional until a cited board replaces them.
  • AA Intelligence Index figures: Artificial Analysis leaderboard, as-of 2026-07-23.

This is a point-in-time snapshot (2026.07.2). Prices and placements move; re-check the linked sources before you commit spend against any single cell.