Platforms

GPU Cloud Platforms in 2026: Three Tiers, Three Tradeoffs

Storage, egress, and idle time routinely beat the hourly rate. Plus why spot capacity is wrong for serving.

Dana Kwon

Contributing Reviewer

Published 6 min read
Three NVIDIA GeForce RTX graphics cards stacked on a surface, showcasing their sleek design and branding details.
Jump to 5 sections

Quick answer: GPU cloud platforms split into three tiers: hyperscalers with enterprise terms and high prices, specialist GPU clouds with better rates and thinner guarantees, and marketplace platforms aggregating spare capacity at the lowest prices with the least reliability. Match the tier to the workload — production inference on the first two, experiments and batch training on the third.

The spread in GPU pricing across platforms is wide enough that the choice matters more than most infrastructure decisions, and the cheapest option is frequently the wrong one for reasons that have nothing to do with the hourly rate.

This roundup covers the GPU cloud landscape in late 2026: the three platform tiers, what you give up at each price point, the commitment structures that change the math, and how to match a tier to a workload.

The three platform tiers

They differ in guarantees more than in silicon, and the guarantees are what you are actually buying at the higher price points.

Two workers handle a package in a spacious warehouse surrounded by shelves stocked with boxes and products.
The three platform tiers
TierRelative priceYou getYou give up
HyperscalersHighestSLAs, regions, compliance, integrationMargin, sometimes availability
Specialist GPU cloudsMiddleBetter rates, newer hardware soonerThinner terms, fewer regions
MarketplacesLowestCheapest capacity availableReliability, support, consistency

Hyperscalers are chosen for procurement reasons at least as often as technical ones. An existing committed spend agreement, an approved data processing addendum, or a compliance requirement frequently decides this before anyone compares hourly rates — the same dynamic that drives cloud selection for inference generally, as we cover in our comparison of inference providers.

Specialist clouds occupy a useful middle. They frequently get new GPU generations before hyperscalers do and price aggressively, at the cost of fewer regions and less contractual depth. For a team that needs current hardware and does not need a signed uptime guarantee, this tier is often the right answer.

Hardware generation matters more than the tier label in one specific case: multi-GPU work. Training and large-model serving depend on interconnect bandwidth between cards, and a platform offering current-generation GPUs over a slow interconnect will underperform an older generation with fast links on exactly the workloads that need several cards. Ask about the interconnect, not just the GPU model — it is the specification most likely to be omitted from a pricing page. The architectural reason is in our comparison of mixture-of-experts and dense models, where expert-parallel serving makes communication overhead a first-order cost.

Marketplaces aggregate spare capacity, including from individual operators. The prices are genuinely low and the variance is genuinely high — in hardware condition, network quality, and whether your instance survives the night.

The costs that are not the hourly rate

Three, and they can exceed the compute line on some workloads.

High angle of shiny wooden ceremonial mallet with golden detail placed on judge tale near documents folders

Storage comes first. Model weights are large — tens to hundreds of gigabytes for bigger checkpoints — and they need to live somewhere fast enough to load quickly. Persistent high-performance storage attached to an idle instance bills continuously, and teams that scale compute to zero often forget that storage did not.

Checkpoint size drives the storage line directly, and it is worth knowing the number before provisioning. Published weights for larger open models run to hundreds of gigabytes, and the file sizes are visible on Hugging Face's model hub before you download anything. A model you intend to keep hot across several instances multiplies that figure by the instance count.

Egress is second and the most punitive. Moving data out of a platform is priced per gigabyte at many providers, which makes multi-cloud architectures and platform migrations expensive in ways that are invisible until the first bill. Check egress pricing before you plan to move anything large. We have seen a team move a 400GB checkpoint off a hyperscaler mid-evaluation and get hit with an egress charge that ran higher than the week of compute they were testing — a five-minute read of the pricing page would have caught it.

Idle time is third and largest in practice. A GPU reserved but unused bills identically to one running flat out, which is the whole argument in our comparison of self-hosting versus API inference. Utilization is the number that decides whether any of this is economical.

Pricing spread across the tiers is wide enough to matter in real dollars, not just in principle. A high-memory GPU instance that runs roughly $2.50 an hour on a hyperscaler frequently lists closer to $1.60 on a specialist cloud and $0.80 to $1.10 on a marketplace for the same card generation, as of late 2026. Over a month of continuous use, that gap is the difference between a four-figure and a low-five-figure bill — which is why the reliability difference has to be worth it before chasing the marketplace number.

How much does GPU cloud actually cost per workload?

A production inference workload running continuously on a mid-tier GPU typically lands between $1,200 and $1,900 a month on a specialist cloud once storage and modest egress are included, versus roughly $1,800 to $2,600 on a hyperscaler for the same hours. Batch and training work on spot capacity can run 60-70% below the on-demand rate, provided the job checkpoints reliably.

Three purchasing modes, and picking the wrong one for a workload is a common and expensive mistake.

Businessperson writes in a planner on a desk, organizing weekly schedule professionally.
  1. On-demand. Highest rate, no commitment, available immediately. Correct for variable or short-lived work.
  2. Reserved or committed. Substantially lower rate in exchange for a term commitment. Correct for steady baseline load you are confident persists.
  3. Spot or preemptible. Lowest rate, with the provider able to reclaim the instance on short notice. Correct for interruptible batch work with checkpointing.

Spot capacity is excellent for training runs and batch inference that checkpoint properly, and it is unsuitable for user-facing serving. An instance reclaimed mid-request is an outage, and building enough redundancy to tolerate that usually costs more than the discount saved.

Availability is the constraint that overrides all three modes during a capacity crunch. The GPU you want at the price you want may simply not be obtainable in your region, and a platform quoting an attractive rate for hardware with a multi-week queue is quoting a theoretical price. Check actual provisioning time before committing to an architecture, and keep a second platform qualified — the multi-provider reasoning in our comparison of inference providers applies with more force to scarce hardware than to API capacity.

Commitments deserve scrutiny in a market moving this fast. A multi-year commitment to a specific GPU generation carries obsolescence risk, and hardware that looks current today may be two generations behind before the term expires. Shorter commitments cost more per hour and preserve optionality that has real value right now.

Run the actual math before signing anything longer than a year. A one-year reserved instance typically discounts 30-40% off on-demand; a three-year term often reaches 50-60% off, which sounds decisive until you price in the risk that the GPU generation is obsolete for your workload well before the term ends. At Model Drop we generally advise treating anything past 18 months as a bet on hardware stability, not just a pricing decision.

Matching a tier to your workload

Four workload shapes cover most of what teams actually run, and each has a clear answer.

Detailed close-up of computer motherboard showing components like RAM slots and capacitors.

Production inference serving users needs availability above all, which means hyperscaler or specialist cloud, on-demand or reserved, with redundancy across at least two instances. This is not the place to optimize the hourly rate.

Batch inference and overnight processing is the ideal spot-capacity workload. It tolerates interruption, saturates the GPU, and runs when demand is low. Checkpoint properly and take the discount.

Fine-tuning and training sits between them. Spot works if your training loop checkpoints reliably; otherwise a reclaimed instance costs more in lost progress than the discount saved.

Experimentation belongs on the cheapest thing available. Marketplace instances are fine when the cost of an interruption is restarting a notebook, and that is most research work.

Cold start time is the variable that determines whether scale-to-zero is viable for any of these. Loading a large checkpoint can take minutes, and a platform that bills for that loading time on every scale-up changes the economics of bursty serving substantially. Measure it before designing around it — the same throughput-verification discipline that MLCommons applies in its MLPerf inference suites applies here, and vendor figures rarely survive contact with a real checkpoint.

The bottom line on GPU cloud platforms

Three tiers selling guarantees rather than silicon. Storage, egress, and idle time frequently exceed the hourly rate in real bills. Spot capacity suits batch and training with checkpointing and is wrong for serving, and long commitments carry real obsolescence risk in a market this fast.

Your next step: price your workload including storage, egress, and realistic idle time before comparing hourly rates. At Model Drop that exercise changes the ranking more often than it confirms it — a platform that looks 20% cheaper on the hourly line has, in our own comparisons, come out more expensive once egress and idle time were added back in often enough that we no longer trust a bare hourly-rate comparison for anything beyond a rough first pass.

Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.

What are the main types of GPU cloud platform?
Three tiers: hyperscalers offering SLAs, broad regions, and compliance at the highest price; specialist GPU clouds with better rates and often newer hardware sooner but thinner terms; and marketplaces aggregating spare capacity at the lowest prices with the least reliability and support.
What costs are hidden beyond the GPU hourly rate?
Persistent storage for large model checkpoints, which bills continuously even when compute scales to zero; egress charges that make migrations and multi-cloud architectures expensive; and idle time, since a reserved GPU bills identically whether it runs flat out or sits unused.
When should I use spot or preemptible GPU instances?
For interruptible batch work that checkpoints reliably — batch inference, overnight processing, and training loops with solid checkpointing. Never for user-facing serving, where an instance reclaimed mid-request is an outage and the redundancy needed to tolerate that costs more than the discount saves.
Are long GPU commitments worth the discount?
They carry real obsolescence risk in a fast-moving hardware market. A multi-year commitment to a specific GPU generation may leave you two generations behind before the term expires. Shorter commitments cost more per hour and preserve optionality that currently has genuine value.
Why does cold start time matter for GPU platforms?
It determines whether scale-to-zero is viable. Loading a large checkpoint can take minutes, and a platform that bills for loading time on every scale-up substantially changes the economics of bursty serving. Measure it on your own checkpoint rather than trusting published figures.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons