GPU Cloud Platforms in 2026: Three Tiers, Three Tradeoffs
Storage, egress, and idle time routinely beat the hourly rate. Plus why spot capacity is wrong for serving.
Jump to 5 sections
Quick answer: GPU cloud platforms split into three tiers: hyperscalers with enterprise terms and high prices, specialist GPU clouds with better rates and thinner guarantees, and marketplace platforms aggregating spare capacity at the lowest prices with the least reliability. Match the tier to the workload — production inference on the first two, experiments and batch training on the third.
The spread in GPU pricing across platforms is wide enough that the choice matters more than most infrastructure decisions, and the cheapest option is frequently the wrong one for reasons that have nothing to do with the hourly rate.
This roundup covers the GPU cloud landscape in late 2026: the three platform tiers, what you give up at each price point, the commitment structures that change the math, and how to match a tier to a workload.
The three platform tiers
They differ in guarantees more than in silicon, and the guarantees are what you are actually buying at the higher price points.
| Tier | Relative price | You get | You give up |
|---|---|---|---|
| Hyperscalers | Highest | SLAs, regions, compliance, integration | Margin, sometimes availability |
| Specialist GPU clouds | Middle | Better rates, newer hardware sooner | Thinner terms, fewer regions |
| Marketplaces | Lowest | Cheapest capacity available | Reliability, support, consistency |
Hyperscalers are chosen for procurement reasons at least as often as technical ones. An existing committed spend agreement, an approved data processing addendum, or a compliance requirement frequently decides this before anyone compares hourly rates — the same dynamic that drives cloud selection for inference generally, as we cover in our comparison of inference providers.
Specialist clouds occupy a useful middle. They frequently get new GPU generations before hyperscalers do and price aggressively, at the cost of fewer regions and less contractual depth. For a team that needs current hardware and does not need a signed uptime guarantee, this tier is often the right answer.
Hardware generation matters more than the tier label in one specific case: multi-GPU work. Training and large-model serving depend on interconnect bandwidth between cards, and a platform offering current-generation GPUs over a slow interconnect will underperform an older generation with fast links on exactly the workloads that need several cards. Ask about the interconnect, not just the GPU model — it is the specification most likely to be omitted from a pricing page. The architectural reason is in our comparison of mixture-of-experts and dense models, where expert-parallel serving makes communication overhead a first-order cost.
Marketplaces aggregate spare capacity, including from individual operators. The prices are genuinely low and the variance is genuinely high — in hardware condition, network quality, and whether your instance survives the night.
The costs that are not the hourly rate
Three, and they can exceed the compute line on some workloads.
Storage comes first. Model weights are large — tens to hundreds of gigabytes for bigger checkpoints — and they need to live somewhere fast enough to load quickly. Persistent high-performance storage attached to an idle instance bills continuously, and teams that scale compute to zero often forget that storage did not.
Checkpoint size drives the storage line directly, and it is worth knowing the number before provisioning. Published weights for larger open models run to hundreds of gigabytes, and the file sizes are visible on Hugging Face's model hub before you download anything. A model you intend to keep hot across several instances multiplies that figure by the instance count.
Egress is second and the most punitive. Moving data out of a platform is priced per gigabyte at many providers, which makes multi-cloud architectures and platform migrations expensive in ways that are invisible until the first bill. Check egress pricing before you plan to move anything large. We have seen a team move a 400GB checkpoint off a hyperscaler mid-evaluation and get hit with an egress charge that ran higher than the week of compute they were testing — a five-minute read of the pricing page would have caught it.
Idle time is third and largest in practice. A GPU reserved but unused bills identically to one running flat out, which is the whole argument in our comparison of self-hosting versus API inference. Utilization is the number that decides whether any of this is economical.
Pricing spread across the tiers is wide enough to matter in real dollars, not just in principle. A high-memory GPU instance that runs roughly $2.50 an hour on a hyperscaler frequently lists closer to $1.60 on a specialist cloud and $0.80 to $1.10 on a marketplace for the same card generation, as of late 2026. Over a month of continuous use, that gap is the difference between a four-figure and a low-five-figure bill — which is why the reliability difference has to be worth it before chasing the marketplace number.
How much does GPU cloud actually cost per workload?
A production inference workload running continuously on a mid-tier GPU typically lands between $1,200 and $1,900 a month on a specialist cloud once storage and modest egress are included, versus roughly $1,800 to $2,600 on a hyperscaler for the same hours. Batch and training work on spot capacity can run 60-70% below the on-demand rate, provided the job checkpoints reliably.
Three purchasing modes, and picking the wrong one for a workload is a common and expensive mistake.
- On-demand. Highest rate, no commitment, available immediately. Correct for variable or short-lived work.
- Reserved or committed. Substantially lower rate in exchange for a term commitment. Correct for steady baseline load you are confident persists.
- Spot or preemptible. Lowest rate, with the provider able to reclaim the instance on short notice. Correct for interruptible batch work with checkpointing.
Spot capacity is excellent for training runs and batch inference that checkpoint properly, and it is unsuitable for user-facing serving. An instance reclaimed mid-request is an outage, and building enough redundancy to tolerate that usually costs more than the discount saved.
Availability is the constraint that overrides all three modes during a capacity crunch. The GPU you want at the price you want may simply not be obtainable in your region, and a platform quoting an attractive rate for hardware with a multi-week queue is quoting a theoretical price. Check actual provisioning time before committing to an architecture, and keep a second platform qualified — the multi-provider reasoning in our comparison of inference providers applies with more force to scarce hardware than to API capacity.
Commitments deserve scrutiny in a market moving this fast. A multi-year commitment to a specific GPU generation carries obsolescence risk, and hardware that looks current today may be two generations behind before the term expires. Shorter commitments cost more per hour and preserve optionality that has real value right now.
Run the actual math before signing anything longer than a year. A one-year reserved instance typically discounts 30-40% off on-demand; a three-year term often reaches 50-60% off, which sounds decisive until you price in the risk that the GPU generation is obsolete for your workload well before the term ends. At Model Drop we generally advise treating anything past 18 months as a bet on hardware stability, not just a pricing decision.
Matching a tier to your workload
Four workload shapes cover most of what teams actually run, and each has a clear answer.
Production inference serving users needs availability above all, which means hyperscaler or specialist cloud, on-demand or reserved, with redundancy across at least two instances. This is not the place to optimize the hourly rate.
Batch inference and overnight processing is the ideal spot-capacity workload. It tolerates interruption, saturates the GPU, and runs when demand is low. Checkpoint properly and take the discount.
Fine-tuning and training sits between them. Spot works if your training loop checkpoints reliably; otherwise a reclaimed instance costs more in lost progress than the discount saved.
Experimentation belongs on the cheapest thing available. Marketplace instances are fine when the cost of an interruption is restarting a notebook, and that is most research work.
Cold start time is the variable that determines whether scale-to-zero is viable for any of these. Loading a large checkpoint can take minutes, and a platform that bills for that loading time on every scale-up changes the economics of bursty serving substantially. Measure it before designing around it — the same throughput-verification discipline that MLCommons applies in its MLPerf inference suites applies here, and vendor figures rarely survive contact with a real checkpoint.
The bottom line on GPU cloud platforms
Three tiers selling guarantees rather than silicon. Storage, egress, and idle time frequently exceed the hourly rate in real bills. Spot capacity suits batch and training with checkpointing and is wrong for serving, and long commitments carry real obsolescence risk in a market this fast.
Your next step: price your workload including storage, egress, and realistic idle time before comparing hourly rates. At Model Drop that exercise changes the ranking more often than it confirms it — a platform that looks 20% cheaper on the hourly line has, in our own comparisons, come out more expensive once egress and idle time were added back in often enough that we no longer trust a bare hourly-rate comparison for anything beyond a rough first pass.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.