Platforms

Inference Providers Compared: Price Is the Wrong First Question

Rate limits and p99 latency decide more deployments than per-token pricing. Plus why identical weights serve differently.

Dana Kwon

Contributing Reviewer

Published 5 min read
Yellow letter tiles spell the word 'price' against a vibrant blue backdrop, ideal for business concepts.
Jump to 5 sections

Quick answer: Inference providers fall into three groups: first-party labs serving their own models, cloud platforms reselling models with enterprise terms, and specialist providers serving open weights at aggressive prices and speeds. Compare on effective cost per completed task, p99 latency, rate limits at your tier, and contractual terms — not on the headline per-token rate.

The per-token price on a provider's landing page is the least informative number in the comparison. Rate limits, cold starts, and tail latency decide far more about whether a provider works for you.

This comparison covers the three provider categories as they stand in late 2026, what each is structurally good at, the four dimensions worth measuring, and how to run a provider evaluation that produces a defensible answer.

The three provider categories

They sell different things, which is why a straight price comparison across them tends to mislead.

Speedometer reading showing speed in km/h on a dark background.
The three provider categories
CategorySellsStrengthWeakness
First-party labsTheir own frontier modelsDay-one access, newest capabilitySingle-vendor exposure
Cloud platformsModels plus enterprise termsProcurement, residency, committed spendLag on new versions, markup
Specialist providersOpen weights, optimized servingPrice and throughputNo frontier closed models

Cloud platforms are frequently chosen for reasons that have nothing to do with the model. Existing committed spend, an established data processing agreement, a region requirement, or a procurement process that already approved the vendor. Those are legitimate and they often outweigh a per-token difference entirely.

The markup for that convenience is usually visible once you compare it directly against the first-party rate for the same model. We have seen it run anywhere from 10% to 40% depending on the platform and the committed-spend tier, which is a real number to weigh against the procurement friction saved rather than an abstract "you pay for the terms" hand-wave.

First-party access has a timing advantage that matters on a fast-moving beat. A model released by its own lab is callable on day one, while the same model reaching a cloud platform's catalog can take weeks — a lag that was visible across the September 2026 release cluster, where several launches were first-party only for an extended window. If being current matters to your product, that lag is a real cost of the cloud route.

Specialist providers compete on serving efficiency over open weights, which is a genuinely different business from training models. The weights are the same ones anyone can download — the product is how fast and cheap they run them.

The four dimensions worth comparing

Price is one of them and not the first.

From below of fiber optic switch with sockets and connected rubber cables on blurred background
  1. Effective cost per completed task. Not per token. Output length varies by model and provider configuration, so measure a real workload end to end.
  2. Rate limits at your account tier. Published separately from pricing and frequently the binding constraint. A cheap provider you cannot call at volume is not cheap.
  3. p99 latency, in your region. Median latency hides the tail, and the tail is what users experience as broken. Cold starts on less-used models are a common cause.
  4. Terms. Data retention, training usage, residency, uptime commitments, and support response. These rarely appear in a comparison table and routinely decide enterprise selections.

Rate limits deserve emphasis because they are the most common surprise. A provider quoting an attractive rate may cap your tier well below the throughput you need, and raising that cap can require a commitment that changes the economics entirely. Check the limit before the price — the same omission pattern our guide to reading launch announcements describes for model releases.

A default tier commonly caps a new account at a few hundred requests per minute, which is comfortable for a demo and tight for anything serving real traffic. Raising it usually means either a sales conversation and a spend commitment, or a multi-week verification process — neither of which shows up anywhere near the pricing page, and both of which should factor into how "cheap" a provider really is for your actual launch timeline.

Why the same weights perform differently

Because serving a model is an engineering problem with a wide quality range, and providers make different tradeoffs on it.

Technician operating laboratory electronic testing and measurement devices with colorful display.

Quantization is the first and least advertised. A provider serving a model at reduced precision delivers more throughput at lower cost and slightly different outputs than the full-precision weights. Providers do not always disclose this, and it is worth asking directly, because it means "the same model" is not always the same model.

The size classes and tradeoffs of quantized serving are covered in more depth in our roundup of running models on a single GPU, but the version relevant here is simpler: a provider quoting the lowest price for a given open-weight model is frequently the one running the most aggressive quantization, and that tradeoff should be an explicit decision, not a discovery made after a quality complaint.

Batching and scheduling policy is the second. Continuous batching and paged attention determine how many concurrent requests a GPU serves efficiently, which sets both cost and latency under load. Standardized measurement of exactly this — under fixed methodology with audited submissions — is what MLCommons publishes in its MLPerf inference suites, and it is better evidence than a provider's own throughput chart.

Hardware and region are the third. The GPU generation behind an endpoint and its physical distance from your users both move latency, and neither is usually visible from the pricing page. A request round-tripping to a data center on another continent can add 100-200ms before the model does any work at all — latency that no amount of serving optimization on the provider's side will claw back.

The practical consequence: benchmark the provider, not the model. Two providers serving identical open weights from the same published checkpoint can differ materially in speed, cost, and output consistency.

How do you actually run a provider evaluation that means something?

Take 100 real requests from your traffic, run them against each candidate provider at realistic concurrency, and record cost, time to first token, latency, and output validity for every request. Compare p99 latency and cost per completed task, not the mean or headline rate. It takes about a day and produces a defensible number.

Top-down view of a laptop keyboard and colorful printed charts on a desk.

The mean will flatter everyone, which is why p99 is the number to defend a decision on. We have run this exact test enough times at Model Drop to see the same pattern repeat: two providers within 5% of each other on mean latency routinely differ by 3-4x at p99, and that tail is what a real user actually experiences during a traffic spike, not the calm average a sales deck shows.

Measure output consistency across providers while you are at it. Sampling parameters, default temperature, and quantization differences mean the same prompt against the same nominal model can produce meaningfully different responses depending on who serves it. If you have prompts tuned against one provider, a migration is not always a configuration change — a version of the portability problem described in our roundup of the model families that matter in late 2026.

Run it at your actual concurrency specifically. Providers that look identical at one request per second diverge sharply at fifty, and the divergence is where the engineering differences live.

Test failure behavior too. Deliberately exceed a rate limit and observe what happens — a clean 429 with a retry-after header is workable, a timeout is not. At Model Drop this is the test that most often changes a provider decision, because it only shows up under stress and never in a sales conversation.

We have watched a provider that looked identical to its competitor on every other dimension get ruled out purely on this test: instead of a clean rejection, it held the connection open for the full timeout window under load, which meant a burst of traffic degraded into a pile of hung requests rather than a clean set of retries. That is the kind of failure a pricing comparison will never surface.

Multi-provider routing has become standard practice rather than an exotic hedge, for availability as much as cost. The architecture is the same one we describe for model tiers in our breakdown of LLM API pricing across the 2026 tiers: default to the cheap path, escalate or fail over on a defined condition.

The bottom line on inference providers

Three categories selling genuinely different things. Effective cost per task, rate limits, p99 latency, and contractual terms are the comparison axes; headline per-token price is not. Identical open weights perform differently across providers because serving is an engineering problem with a wide quality range.

Your next step: run 100 of your own requests against two providers at realistic concurrency and compare p99 latency and cost per completed task. A day of work settles a decision that a pricing page cannot.

Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.

What types of inference providers are there?
Three: first-party labs serving their own frontier models, cloud platforms reselling models alongside enterprise terms and procurement, and specialist providers serving open weights with optimized infrastructure. They sell genuinely different things, which is why comparing them on per-token price alone tends to mislead.
Why do two providers serving the same model perform differently?
Serving is an engineering problem with a wide quality range. Providers make different choices on quantization precision, batching and scheduling policy, GPU generation, and region placement. A provider serving reduced-precision weights delivers more throughput and slightly different outputs, and this is not always disclosed.
What should I compare besides price?
Effective cost per completed task rather than per token, rate limits at your specific account tier, p99 latency in your region, and contractual terms covering retention, training usage, residency, and uptime. Rate limits are the most common surprise, since a cheap provider you cannot call at volume is not cheap.
Why does p99 latency matter more than median?
Median latency hides the tail, and the tail is what users experience as broken. Cold starts on less-frequently-used models are a common cause of bad p99 numbers that never appear in average-case marketing figures or in a light evaluation run at low concurrency.
Should I use more than one inference provider?
Multi-provider routing has become standard practice, for availability as much as cost. The pattern is to default to a cheap path and escalate or fail over on a defined condition. Test failure behavior explicitly: a clean 429 with retry-after is workable, a timeout is not.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons