LLM API Pricing in 2026: The Three-Tier Ladder Every Lab Runs
Frontier to budget spans three orders of magnitude. Picking the right rung matters more than picking the right vendor.
Jump to 7 sections
Quick answer: LLM API pricing in 2026 spans roughly three orders of magnitude, from about $10 per million input tokens at the frontier tier down to a few cents for small open-weight models. Every major lab now runs the same three-tier ladder, and picking the right rung matters more than picking the vendor.
The interesting thing about LLM API pricing is no longer which lab is cheapest. It is that every lab has converged on the same shape, and most teams are running on the wrong rung of it.
This roundup covers how LLM API pricing works in 2026: the tier structure every vendor uses, what the frontier and budget ends actually cost, the discounts that change the math, and how to price a real workload rather than comparing headline rates. It is written for whoever signs off on the inference bill.
Every vendor now runs the same three-tier ladder
Flagship, workhorse, cheap. That is the structure, and it is remarkably consistent across labs that otherwise agree on very little.
The flagship tier sits around $10 per million input tokens and carries an output price four to five times that. The workhorse tier lands near $2 to $4 input. The cheap tier drops to cents. Each lab uses different names for these rungs, which obscures how similar the ladders are.
Anthropic's published pricing is the clearest published example of the pattern, listing Fable 5.1 at $10 input and $50 output per million tokens, Opus 5.5 at $4 and $20, Sonnet 5 at $2 and $10, and Haiku 4.5 at $1 and $5. That is four rungs rather than three, and the spread from top to bottom is tenfold.
OpenAI's ladder has the same slope, with the GPT-6 line splitting into a top tier near $10 input and substantially cheaper variants beneath it; check OpenAI's platform pricing documentation for current figures, since promotional rates and tier names change more often than the structure does. Google runs a Pro and Flash split on the same logic.
What the frontier tier actually costs to run
More than the input price suggests, because output dominates. At $50 per million output tokens, an agent run that emits 40,000 tokens costs $2.00 in output alone, before a single input token is counted.
| Workload | Output tokens | At $50/1M | At $10/1M | At $0.50/1M |
|---|---|---|---|---|
| Short classification | 200 | $0.01 | $0.002 | $0.0001 |
| Document summary | 1,500 | $0.075 | $0.015 | $0.0008 |
| Long agent run | 40,000 | $2.00 | $0.40 | $0.02 |
| 10k runs/month | 40,000 each | $20,000 | $4,000 | $200 |
That bottom row is the one that ends arguments. The same monthly volume costs $20,000 or $200 depending purely on which rung you picked, and for a large share of production traffic the cheap rung produces an indistinguishable result.
Reasoning-heavy models complicate this further. A model that thinks longer emits more output for the same nominal task, so its effective cost per completed job can exceed a nominally pricier model that answers directly. Compare cost per task, never cost per token.
The budget end has gotten genuinely cheap
Open-weight and small hosted models now sit two to three orders of magnitude below the frontier. Several providers list sub-cent input pricing per million tokens for small models, and cache-hit rates on some APIs drop the effective input cost further still.
The capability gap has narrowed enough to matter. Open-weight families published on Hugging Face's model hub now cover most routine production work — classification, extraction, summarization, routing — at a fraction of frontier cost. They are not frontier replacements for hard reasoning, and they do not need to be.
At Model Drop we would frame the decision this way: the frontier tier is for the 5% of requests where a wrong answer is expensive. Everything else belongs somewhere cheaper, and the engineering work is in routing between them rather than in picking one.
Caching and batching change the real rate
The list price is rarely what you pay. Two standard discounts move the number substantially, and both require you to structure requests deliberately.
Prompt caching is the big one. Repeated prefixes — a system prompt, a tool schema, a long document you query several times — can be cached, and cache reads are priced far below fresh input tokens. For any workload with a stable preamble, this is usually the single largest available saving.
Caching has a shape requirement worth understanding. The cached portion has to be a stable prefix, so anything that varies — a timestamp, a user id, a retrieved document — must come after the static part of the prompt. Teams that interleave dynamic content into their system prompt get no cache hits and never find out why. Reordering the prompt is usually a one-line change with a large bill impact.
Batch processing is the other. Most vendors offer roughly half price for asynchronous jobs with a delayed completion window. If your workload is a nightly run rather than a user-facing request, you are leaving money on the table by not using it.
Watch the fine print: introductory and promotional rates carry end dates. A tier priced attractively through a stated month reverts afterward, and budgets built on the promotional number break quietly.
How do you price a real workload instead of a rate card?
Measure, then multiply. Take 100 representative requests, log actual input and output token counts, and compute the mean cost per request at each candidate tier. Multiply that by projected monthly volume. This takes about an afternoon and is more reliable than any rate-card comparison, including this one.
Do not skip the tokenizer question either. Token counts differ between vendors for identical text, sometimes by 10% or more, because each lab uses its own tokenizer. A per-token price that looks 5% cheaper can be more expensive once the same document tokenizes differently. Both major labs publish token-counting endpoints, and OpenAI's public tokenizer tool is the quickest way to see the effect on your own text.
Three factors distort estimates built from rate cards alone. Output length varies far more than input length. Retries and tool-calling loops multiply token counts in ways a single-shot estimate misses. And the token-per-minute rate limit on your account tier can force architectural changes that cost more than the tokens do.
Once you have real numbers, the routing question answers itself. The methodology for judging whether a new tier is worth testing is the same one we lay out in our guide to reading a model launch announcement: price the workload before you believe the chart.
Rate limits are part of the real price
A cheaper per-token rate does not always translate into a cheaper deployment, because every vendor also caps tokens-per-minute and requests-per-minute per account tier. A workload that would clear a mid-tier rate card can still get throttled if the account's rate limit hasn't scaled with usage.
Anthropic, OpenAI, and Google all gate higher rate-limit tiers behind cumulative spend or a manual increase request, not just a credit card on file. A team that switches to a cheaper model but keeps hitting 429 errors ends up building retry logic and backoff queues, and that engineering time is a real cost the rate card never shows.
The practical fix is to request a rate-limit increase before a migration, not after the first production incident. Most vendors process these requests within one to three business days when usage history supports the ask, and doing it ahead of a rollout avoids a launch-week scramble that costs more in engineering hours than the token savings were worth that month.
This also affects the caching and batching discounts covered above: a batch job with a tight rate-limit ceiling takes longer to clear, which can push a nightly job past its window. Check the requests-per-minute ceiling for the batch endpoint specifically — it is often lower than the standard endpoint's limit, and vendors document it separately.
The bottom line on LLM API pricing in 2026
Pricing has converged on a three-to-four rung ladder at every major lab, spanning roughly $10 per million input tokens at the top to cents at the bottom. Output dominates the bill, caching and batching cut it materially, and the rung you choose matters far more than the vendor you choose.
Your next step: log token counts on 100 real requests and compute your actual cost per task. Most teams discover they are running frontier-tier inference on workloads that a mid-tier model handles identically.
By Derek Plummer, Staff Writer at Model Drop. Prices cited are vendor list rates as of September 2026 and change frequently — verify against the vendor's own pricing page before budgeting.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.