LLM pricing and capability matrix: the provider table I route agents against

A flat editorial illustration of a routing switchboard with labelled lanes fanning out to four model provider cards, each card showing a price tag and a context-window gauge.

A maintained provider matrix for routing AI agents: context window, tool use, price per 1M tokens, latency, and MCP support, with every figure sourced to the provider's own docs and a published methodology so anyone can re-run it.

TLDR

This is a maintained provider matrix for routing agents: context window, tool use, price per 1M tokens, latency, and MCP support, with every number traced to the provider's own documentation on the date it was checked. The headline finding is that a price per million tokens is not a comparable unit across providers, and in one documented case a 1 percent increase in prompt size raises the bill for a single run by 82 percent. Cells that could not be sourced say "not published" instead of guessing.

By the founder of Cerevisor. All figures below were checked against provider documentation on 24 July 2026. That date is the whole point of the page, so it appears again in the methodology.

Here is the moment that made me build this table. A workflow was running a long research agent on Gemini 3.1 Pro Preview. Prompt size drifted from roughly 199,000 tokens to roughly 201,000 tokens as the run picked up more files. Two thousand extra tokens. One percent. The cost of that run went from about $0.64 to about $1.16, because Google’s published pricing doubles the input rate and lifts output by half above a 200,000 token prompt. Nothing broke. No error fired. The bill simply grew 82 percent while the work stayed the same.

That is not a pricing gotcha. That is a routing input, and it belongs in a table next to context window and tool support.


Why an LLM pricing comparison stops being apples to apples

Every provider publishes dollars per million tokens. The trap is that the token is not the same unit on both sides of the comparison.

Anthropic says so directly in its own pricing documentation:

"Claude Opus 4.7 and later Opus models, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, and Claude Sonnet 5 use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text."

Anthropic Claude Docs, pricing page, checked July 2026

Read that against a rate card and the arithmetic goes sideways. Two models can show the same dollars per million and still bill differently for the identical prompt, because one counts the prompt into more tokens than the other. Any LLM pricing comparison that stops at the rate card is comparing labels, not spend.

Two more unit problems sit in the same bucket. Google’s pricing page states that thinking tokens are included in output pricing, so a model that reasons more before answering costs more per finished answer at an unchanged output rate. And cache reads land at a tenth of the input rate on both Anthropic and OpenAI, which means a well-cached agent and a badly cached agent on the same model are effectively on different price plans.

The clearest live example landed three days before this table was compiled. Google shipped Gemini 3.6 Flash on 21 July 2026 at $1.50 input and $7.50 output, against 3.5 Flash at the same $1.50 input and $9.00 output. The rate card shows one discount, 17 percent off the output line. Google’s announcement describes a second one:

"on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash"

Google, Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, 21 July 2026. The same post reports up to 65 percent fewer on DeepSWE by Datacurve.

Those two discounts compound, and only one of them is printed on a pricing page. Apply the documented 17 percent to the reference run below and the arithmetic moves: a task that spends 20,000 output tokens on 3.5 Flash should land near 16,600 on 3.6 Flash, which prices the run at roughly $0.42 rather than the $0.45 the rate card implies, against $0.48 for its predecessor. That is a computed illustration of a published figure, not a measurement of your workload, and it is the whole argument in miniature. A cheaper rate and a quieter model are different savings, and a table of rates can only see the first.

Key Insight

A rate card prices a token. A routing decision prices a run. The two only agree when the tokenizer, the thinking tokens, and the cache behavior are held constant, which across providers they never are.


The provider matrix: context, tool use, price per 1M, and MCP support

Capability axes first. Every cell here is either quoted from provider documentation or marked as unavailable.

Provider capability matrix, documentation checked 24 July 2026
ProviderContext windowMax outputFunction callingRemote MCPLatency
Anthropic Claude API1M (Opus 4.8, Sonnet 5, Fable 5); 200k (Haiku 4.5)128k (Haiku 4.5: 64k)Yes, with published token overhead per modelYes, Messages API connector, beta header mcp-client-2025-11-20; tool calls only; HTTPS remote servers only, no local stdioComparative class only: Fastest, Fast, Moderate, Slower. No absolute figures published
OpenAI GPT-5.6 (Sol, Terra, Luna)1.05M128kYes (functions, web search, file search, computer use)Yes, as a tool type in the Responses API; Streamable HTTP or HTTP/SSE; require_approval and allowed_tools controlsNot published
Google Gemini API (current lineup: 3.6 Flash, 3.5 Flash-Lite, 3.1 Pro Preview; no 3.6 Pro is published)Not published on the models overview page; exposed programmatically as inputTokenLimitNot published on that page (outputTokenLimit via API)Yes, four modes: auto, any, none, validated (preview)Yes, via the Interactions API; Streamable HTTP only, not SSENot published
DeepSeek (OpenAI-compatible)1M384kYes, through the OpenAI-compatible surfaceNot publishedNot published
Local, OpenAI-compatible (Ollama, vLLM)Set by the model and the serverSet by the serverModel dependentClient side onlyNot measured here; it is a property of the machine, so measure it on the machine

Now the money. These are list rates from each provider’s published pricing page on 24 July 2026, in USD per 1M tokens.

Published rates per 1M tokens, and one reference run
ModelInputOutputReference run (200k in, 20k out, no cache)
gpt-5.6-sol$5.00$30.00$1.60
claude-opus-4-8$5.00$25.00$1.50
gpt-5.6-terra$2.50$15.00$0.80
gemini-3.1-pro-preview (prompt under 200k)$2.00$12.00$0.64
claude-sonnet-5 (intro rate to 31 Aug 2026)$2.00$10.00$0.60
gemini-3.5-flash$1.50$9.00$0.48
gemini-3.6-flash (released 21 Jul 2026)$1.50$7.50$0.45
gpt-5.6-luna$1.00$6.00$0.32
claude-haiku-4-5$1.00$5.00$0.30
gemini-3.5-flash-lite$0.30$2.50$0.11
deepseek-v4-pro$0.435 (cache miss)$0.87$0.104
deepseek-v4-flash$0.14 (cache miss)$0.28$0.034
47x
spread between the most and least expensive published option for the same 200k-in, 20k-out reference run

A few rows carry expiry dates rather than prices. Claude Sonnet 5 lists $2 and $10 as introductory pricing through 31 August 2026, reverting to $3 and $15 on 1 September. DeepSeek’s documentation notes that the deepseek-chat and deepseek-reasoner model names are deprecated on 2026/07/24, which is the day this table was compiled. A matrix without a measurement date is a matrix that lies quietly.

The Gemini rows are three days old at time of writing. Google’s 21 July 2026 release covered 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, the last of which is a security-tuned model restricted to a limited-access pilot for governments and trusted partners, so it carries no public rate and gets no row here. 3.5 Flash stays in the table as the predecessor, because the comparison against it is where the token-efficiency argument shows up. There is still no Gemini 3.6 Pro on the pricing page or the models page, which means the frontier Gemini slot for a long-context routing decision remains 3.1 Pro Preview and its 200,000 token cliff.


How this LLM API pricing comparison was measured

The methodology is short on purpose, because a model comparison chart nobody can reproduce is a marketing asset, not a reference.

Documented axes. Context window, max output, function calling, remote MCP support, and price per 1M tokens come from the provider’s own pricing or model documentation. Every source URL is listed at the end of this article, and each was fetched and read on 24 July 2026. Where a page does not state a value, the cell says “not published” rather than borrowing a number from a third-party aggregator.

Computed axis. The reference run column is arithmetic, not a benchmark: 200,000 input tokens and 20,000 output tokens at the published rates, no caching, no batch discount. Anyone can recompute it in a spreadsheet in a minute. It exists because rate cards mislead about relative order once input-to-output ratios differ.

Not measured. Latency. I did not run a timed test across these providers for this edition, so there are no latency numbers here beyond the comparative class Anthropic publishes for its own models. Inventing them would have been easy and would have destroyed the point of the page. The procedure for the next edition is written down and simple: fifty identical requests per model at a fixed prompt size, streaming on, recording time to first token and total wall time, from one machine on one network, discarding the first five as warm-up, and publishing median with the interquartile range. Until that runs, the cell says “not published”.

Refresh cadence. Quarterly, or sooner when a provider changes a rate. The intro-price expiry on 31 August 2026 is already on the calendar.


The cheapest LLM API is not the cheapest agent run

Three published details break the ranking above, and each one is a routing rule rather than trivia.

The tier cliff comes first. Gemini 3.1 Pro Preview prices prompts at or under 200,000 tokens at $2 input and $12 output, and above that boundary at $4 and $18. That is the 82 percent jump from the opening. Cerevisor’s provider layer resolves which side of the cliff a request landed on by reading usageMetadata.promptTokenCount off the response, because there is no reliable way to know beforehand once an agent is assembling its own context.

Caching comes second. A cache read costs a tenth of the input rate on Anthropic and on OpenAI’s cached input line. On the reference run, moving 150,000 of the 200,000 input tokens into cache hits takes an Opus 4.8 run from $1.50 to about $0.83. The cheapest model on the rate card can lose to a more expensive one that caches well.

Tool overhead comes third, and only one provider publishes it. Anthropic documents the token cost of the tool-use system prompt per model: 290 tokens for Opus 4.8 with tool choice auto, 675 for Opus 4.7, 496 for Haiku 4.5. For OpenAI and Google, that overhead is real but not published, so it cannot be modelled and has to be observed in your own usage data.

The cheapest LLM API on the rate card and the cheapest provider for an agent that holds a long context and calls tools every turn are frequently not the same product.


Where provider routing actually happens in a workflow

A matrix is only useful if something consumes it. In Cerevisor, provider choice resolves per agent through a four-level chain: an agent-level override wins, then a workflow-level override, then the library default, then the first enabled credential. That sounds bureaucratic until the first time a workflow needs two different providers in the same run, which happens almost immediately.

The four patterns this table feeds, straight out of the per-agent provider overrides guide:

  • A senior research agent on a frontier model, with its reviewer on the fast cheap tier. On the reference run that pairing costs $1.50 plus $0.30 instead of $3.00, and the reviewer’s job does not need frontier reasoning.
  • Sensitive data agents pinned to a local OpenAI-compatible endpoint so the payload never leaves the machine, while public-data agents run on a hosted API. The price column for the local row is not a number, it is a hardware bill.
  • Long-running agents parked on whichever provider survives a laptop sleep, which is a reliability axis rather than a cost axis.
  • MCP-dependent agents routed only to providers whose remote MCP support matches the server in question. Anthropic’s connector handles tool calls over HTTPS and cannot reach a local stdio server. Google’s route accepts Streamable HTTP but not SSE. That difference decides which agent can hold which MCP server, regardless of price.

Three billing failures this table exists to prevent

These are failures I hit in the harness’s own cost tracking, with the fix that closed each one.

A prefix match that overbilled by 5x. Model pricing is resolved by longest-prefix match, so gemini-3.5-flash-lite matched the gemini-3.5-flash row when the more specific row was missing. Flash-Lite bills $0.30 input and $2.50 output; Flash bills $1.50 and $9.00. Every Flash-Lite call was reported at five times the real input cost and 3.6 times the real output cost. The fix was a specific row plus a test asserting that the resolver never falls through to a shorter key when a longer one exists.

A silent zero. An unknown model on an OpenAI-compatible endpoint estimated cost as $0. For someone on local Ollama that is correct and welcome. For someone on a paid aggregator it means a genuinely expensive route looked free on the dashboard. The fix was a one-time warning per model, surfaced as a visible notification rather than a console line nobody reads.

A backoff that ignored the server. Transport errors were collapsed into a plain string error before the retry helper saw them, so a provider replying Retry-After: 30 was unreachable and the loop burned its attempts on 500ms, 1s, and 2s waits before failing anyway. The fix was a typed error that carries the HTTP status and the parsed retry delay, so the backoff honours what the provider actually asked for.

Key Insight

Cost tracking fails quietly. Nothing crashes when a model is billed at the wrong rate, which is exactly why the fallback policy matters more than the table: an unknown model should round against you, not in your favour.


What to check before your next provider swap

  1. Write down the token unit before comparing rates

    Note which models changed tokenizer, and which count thinking tokens as output. A 30 percent tokenizer shift outweighs most rate differences between adjacent tiers.

  2. Price a reference run, not a rate

    Pick one real workload shape from a workflow that already runs, fix the input and output token counts, and price every candidate against it. The ordering will not match the rate card.

  3. Find the tier cliffs and the expiry dates

    Context-tier pricing and introductory rates are both live in this table right now. Put the 31 August 2026 Sonnet 5 revert in a calendar with a named owner.

  4. Make unknown models loud

    Decide what an unrecognised model costs in your accounting, choose the conservative direction, and make sure someone sees a notification the first time it happens.

  5. Route per agent, then re-measure

    Move one agent, not the whole workflow, then compare the actual token counts before and after. Open the provider overview guide and set a single override on the cheapest-to-move agent in the graph.

The concrete next action is the smallest one on that list: pick the one agent in your busiest workflow whose output nobody would notice moving down a tier, route just that agent, and compare a week of token counts. That single override usually pays for the hour it takes to read this table properly.

This page will be re-measured quarterly, and sooner if a provider moves a rate. If a cell here says “not published” and you have a documented source that fills it, that is the most useful thing anyone can send me.

Sources

  1. Pricing - Anthropic Claude Docs, 2026-07-24
  2. Models overview - Anthropic Claude Docs, 2026-07-24
  3. MCP connector - Anthropic Claude Docs, 2026-07-24
  4. API pricing - OpenAI Developer Docs, 2026-07-24
  5. Connectors and remote MCP servers - OpenAI Developer Docs, 2026-07-24
  6. Gemini API pricing - Google AI for Developers, 2026-07-24
  7. Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber - Google, 2026-07-21
  8. Function calling with the Gemini API - Google AI for Developers, 2026-07-24
  9. Models and pricing - DeepSeek API Docs, 2026-07-24

Back to all insights