Inference Providers Compared: Price Is the Wrong First Question
Rate limits and p99 latency decide more deployments than per-token pricing. Plus why identical weights serve differently.
Jump to 5 sections
Quick answer: Inference providers fall into three groups: first-party labs serving their own models, cloud platforms reselling models with enterprise terms, and specialist providers serving open weights at aggressive prices and speeds. Compare on effective cost per completed task, p99 latency, rate limits at your tier, and contractual terms — not on the headline per-token rate.
The per-token price on a provider's landing page is the least informative number in the comparison. Rate limits, cold starts, and tail latency decide far more about whether a provider works for you.
This comparison covers the three provider categories as they stand in late 2026, what each is structurally good at, the four dimensions worth measuring, and how to run a provider evaluation that produces a defensible answer.
The three provider categories
They sell different things, which is why a straight price comparison across them tends to mislead.
| Category | Sells | Strength | Weakness |
|---|---|---|---|
| First-party labs | Their own frontier models | Day-one access, newest capability | Single-vendor exposure |
| Cloud platforms | Models plus enterprise terms | Procurement, residency, committed spend | Lag on new versions, markup |
| Specialist providers | Open weights, optimized serving | Price and throughput | No frontier closed models |
Cloud platforms are frequently chosen for reasons that have nothing to do with the model. Existing committed spend, an established data processing agreement, a region requirement, or a procurement process that already approved the vendor. Those are legitimate and they often outweigh a per-token difference entirely.
The markup for that convenience is usually visible once you compare it directly against the first-party rate for the same model. We have seen it run anywhere from 10% to 40% depending on the platform and the committed-spend tier, which is a real number to weigh against the procurement friction saved rather than an abstract "you pay for the terms" hand-wave.
First-party access has a timing advantage that matters on a fast-moving beat. A model released by its own lab is callable on day one, while the same model reaching a cloud platform's catalog can take weeks — a lag that was visible across the September 2026 release cluster, where several launches were first-party only for an extended window. If being current matters to your product, that lag is a real cost of the cloud route.
Specialist providers compete on serving efficiency over open weights, which is a genuinely different business from training models. The weights are the same ones anyone can download — the product is how fast and cheap they run them.
The four dimensions worth comparing
Price is one of them and not the first.
- Effective cost per completed task. Not per token. Output length varies by model and provider configuration, so measure a real workload end to end.
- Rate limits at your account tier. Published separately from pricing and frequently the binding constraint. A cheap provider you cannot call at volume is not cheap.
- p99 latency, in your region. Median latency hides the tail, and the tail is what users experience as broken. Cold starts on less-used models are a common cause.
- Terms. Data retention, training usage, residency, uptime commitments, and support response. These rarely appear in a comparison table and routinely decide enterprise selections.
Rate limits deserve emphasis because they are the most common surprise. A provider quoting an attractive rate may cap your tier well below the throughput you need, and raising that cap can require a commitment that changes the economics entirely. Check the limit before the price — the same omission pattern our guide to reading launch announcements describes for model releases.
A default tier commonly caps a new account at a few hundred requests per minute, which is comfortable for a demo and tight for anything serving real traffic. Raising it usually means either a sales conversation and a spend commitment, or a multi-week verification process — neither of which shows up anywhere near the pricing page, and both of which should factor into how "cheap" a provider really is for your actual launch timeline.
Why the same weights perform differently
Because serving a model is an engineering problem with a wide quality range, and providers make different tradeoffs on it.
Quantization is the first and least advertised. A provider serving a model at reduced precision delivers more throughput at lower cost and slightly different outputs than the full-precision weights. Providers do not always disclose this, and it is worth asking directly, because it means "the same model" is not always the same model.
The size classes and tradeoffs of quantized serving are covered in more depth in our roundup of running models on a single GPU, but the version relevant here is simpler: a provider quoting the lowest price for a given open-weight model is frequently the one running the most aggressive quantization, and that tradeoff should be an explicit decision, not a discovery made after a quality complaint.
Batching and scheduling policy is the second. Continuous batching and paged attention determine how many concurrent requests a GPU serves efficiently, which sets both cost and latency under load. Standardized measurement of exactly this — under fixed methodology with audited submissions — is what MLCommons publishes in its MLPerf inference suites, and it is better evidence than a provider's own throughput chart.
Hardware and region are the third. The GPU generation behind an endpoint and its physical distance from your users both move latency, and neither is usually visible from the pricing page. A request round-tripping to a data center on another continent can add 100-200ms before the model does any work at all — latency that no amount of serving optimization on the provider's side will claw back.
The practical consequence: benchmark the provider, not the model. Two providers serving identical open weights from the same published checkpoint can differ materially in speed, cost, and output consistency.
How do you actually run a provider evaluation that means something?
Take 100 real requests from your traffic, run them against each candidate provider at realistic concurrency, and record cost, time to first token, latency, and output validity for every request. Compare p99 latency and cost per completed task, not the mean or headline rate. It takes about a day and produces a defensible number.
The mean will flatter everyone, which is why p99 is the number to defend a decision on. We have run this exact test enough times at Model Drop to see the same pattern repeat: two providers within 5% of each other on mean latency routinely differ by 3-4x at p99, and that tail is what a real user actually experiences during a traffic spike, not the calm average a sales deck shows.
Measure output consistency across providers while you are at it. Sampling parameters, default temperature, and quantization differences mean the same prompt against the same nominal model can produce meaningfully different responses depending on who serves it. If you have prompts tuned against one provider, a migration is not always a configuration change — a version of the portability problem described in our roundup of the model families that matter in late 2026.
Run it at your actual concurrency specifically. Providers that look identical at one request per second diverge sharply at fifty, and the divergence is where the engineering differences live.
Test failure behavior too. Deliberately exceed a rate limit and observe what happens — a clean 429 with a retry-after header is workable, a timeout is not. At Model Drop this is the test that most often changes a provider decision, because it only shows up under stress and never in a sales conversation.
We have watched a provider that looked identical to its competitor on every other dimension get ruled out purely on this test: instead of a clean rejection, it held the connection open for the full timeout window under load, which meant a burst of traffic degraded into a pile of hung requests rather than a clean set of retries. That is the kind of failure a pricing comparison will never surface.
Multi-provider routing has become standard practice rather than an exotic hedge, for availability as much as cost. The architecture is the same one we describe for model tiers in our breakdown of LLM API pricing across the 2026 tiers: default to the cheap path, escalate or fail over on a defined condition.
The bottom line on inference providers
Three categories selling genuinely different things. Effective cost per task, rate limits, p99 latency, and contractual terms are the comparison axes; headline per-token price is not. Identical open weights perform differently across providers because serving is an engineering problem with a wide quality range.
Your next step: run 100 of your own requests against two providers at realistic concurrency and compare p99 latency and cost per completed task. A day of work settles a decision that a pricing page cannot.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.