Self-Hosting vs API Inference: Utilization Decides It
Free weights are not free inference. The cost case lives on a utilization curve, and residency is a better reason anyway.
Jump to 5 sections
Quick answer: Self-hosting beats an API only above sustained GPU utilization, typically well past 40% to 50% of a reserved card's capacity running around the clock. Below that, idle silicon costs more than the tokens would have. Data residency, version pinning, and fine-tuning are better reasons to self-host than cost is.
The cost case for self-hosting is made with a per-token comparison and lost on a utilization curve. Free weights are not free inference, and the gap between those two facts is where a lot of infrastructure budget goes to die.
This comparison covers self-hosting versus API inference honestly: the cost arithmetic including the parts people leave out, the three non-cost reasons that actually justify it, and the hybrid arrangement most teams land on.
The cost arithmetic, done properly
A reserved GPU costs the same whether it processes a million tokens an hour or none. That single fact governs the whole comparison, and it is why utilization is the variable to solve for.
| Utilization | Effective cost per token | Versus API |
|---|---|---|
| Under 20% | Very high | API wins clearly |
| 30-50% | Comparable | Roughly a wash |
| 60-80% | Low | Self-hosting wins |
| Sustained near capacity | Lowest available | Self-hosting wins decisively |
The exact crossover depends on your hardware cost, the model, and the API rate you are comparing against. The shape does not change: the curve is steep at low utilization and flattens once the card is busy.
To put a real number on it: a reserved 24 GB GPU running around $1,500-2,000 a month against an API charging $10 per million output tokens crosses over somewhere near 50,000-80,000 requests a day of moderate length, depending on the model. Below that volume, the API is usually cheaper once every hidden cost is counted; above it, owned hardware starts to win decisively.
Traffic shape therefore matters more than traffic volume. A workload averaging 30% utilization because it runs hard for eight hours and idles for sixteen is not a 30% workload for costing purposes — it is a peak-capacity workload with a long idle tail, and you provision for the peak.
Throughput assumptions are where these models most often go wrong. A cost estimate built on a vendor's advertised tokens per second overstates what you will achieve, because published figures come from tuned configurations at ideal batch sizes. Standardized inference measurement under fixed methodology with audited submissions — what MLCommons publishes in its MLPerf inference suites — gives a far more defensible planning number, and even that should be verified on your own hardware and prompt distribution.
Batch workloads are the clean win. Overnight processing, backfills, and bulk enrichment can be scheduled to saturate a GPU, which is exactly the condition self-hosting needs. User-facing traffic with a daytime peak is the clean loss.
The costs that get left out
Three of them, and together they frequently exceed the GPU bill.
Engineering time is the largest. Standing up a serving stack, configuring batching, handling model loading and failover, monitoring throughput, and keeping the whole thing patched is ongoing work by people who cost more per hour than the hardware does. Teams price the GPU and forget the engineer.
We have seen this initial setup alone run two to four weeks of a mid-level engineer's time before the first production request is served, and ongoing maintenance — patching, monitoring, upgrading the serving stack as new quantization formats ship — typically consumes several hours a week indefinitely afterward. At a fully-loaded engineering cost, that ongoing maintenance alone can exceed the GPU's own monthly bill.
Redundancy is the second. A single GPU is a single point of failure, and a production deployment that cannot tolerate one card going down needs at least two — which roughly halves your effective utilization on the same traffic.
That doubling is easy to forget when the original sizing exercise only ever priced one card. A team that sized for 70% utilization on a single GPU and then added a second for redundancy is really running at roughly 35% blended utilization across the pair — which, per the earlier table, is exactly the range where the cost comparison against an API turns from "clearly favorable" to "roughly a wash."
Capacity headroom is the third. Provisioning for peak means paying for peak continuously, and the gap between average and peak is pure cost. Elastic API pricing absorbs that variance for you, which is a real service even when the per-token rate looks worse.
Against those, the token price differential can look smaller than expected — the ladder it is being compared against is in our breakdown of LLM API pricing across the 2026 tiers, where the cheap band has fallen far enough to undercut a poorly-utilized GPU.
What actually justifies self-hosting, if not cost?
Three reasons justify self-hosting on their own, and none of them is cost: data residency requirements that no API arrangement satisfies, version pinning so a model cannot be deprecated or silently changed on someone else's schedule, and fine-tuning on proprietary data you cannot transmit elsewhere. Any one of these alone is sufficient reason to self-host regardless of the cost comparison.
- Data residency. Inference inside your own network boundary or a specific jurisdiction, with no prompt content leaving. For regulated workloads this is binary, and no API arrangement satisfies it.
- Version pinning. A model you control cannot be deprecated on someone else's schedule, and no floating alias can silently repoint it. For systems requiring multi-year behavioral stability, this is decisive.
- Fine-tuning on proprietary data. Training on data you cannot transmit, and keeping the resulting weights. A tuned small model frequently beats a much larger general model on the specific task it was tuned for.
That third reason changes the capability comparison rather than the cost one, and it is the most underused. The economics of running tuned small models are covered in our roundup of models that run on a single GPU.
Licensing is the precondition for all three, and it is checked less often than it should be. Open weights do not uniformly permit every commercial deployment — terms range from permissive MIT and Apache 2.0 to custom licenses with acceptable-use clauses or scale thresholds. The authoritative text sits in the repository on Hugging Face's model hub rather than in any announcement, and it is worth reading before infrastructure is provisioned rather than after.
If one of these applies, run the deployment and treat cost as a secondary optimization. If none applies, cost is the whole argument and it is usually weaker than it looks.
The arrangement most teams end up with
Hosted open-weight endpoints for bulk traffic, a frontier API for hard requests, and self-hosting only where a residency requirement forces it.
The middle option gets skipped in most self-hosting discussions and deserves not to be. A hosted endpoint serving open weights captures most of the cost advantage — cents per million tokens rather than dollars — with none of the operational burden, and it preserves the option to move the same weights in-house later. The tradeoffs are laid out in our comparison of open-weight and closed models.
Where self-hosting does run, sizing it for steady baseline load and bursting to an API for peaks is the configuration that holds up. You get high utilization on the hardware you own and elastic capacity for the variance, without provisioning for a peak you hit twice a day.
At Model Drop our observation is that teams which self-host for cost reasons alone frequently migrate back within a year, while teams that self-host for residency or version-stability reasons stay. The requirement outlasts the spreadsheet.
We have watched that migration-back happen in a fairly consistent shape: a team self-hosts expecting a clean cost win, discovers the engineering and redundancy overhead six months in, and quietly moves back to an API once the on-call burden of running inference infrastructure outweighs whatever margin the spreadsheet originally promised. The teams that stay are the ones who never made cost the primary justification in the first place.
The bottom line
Self-hosting wins on cost only at sustained high utilization, and engineering time, redundancy, and peak headroom usually erase a thin margin. Residency, version pinning, and proprietary fine-tuning justify it regardless of cost. The hosted open-weight middle path is the option most comparisons skip and most teams should try first.
Your next step: plot your actual hourly request volume for a full week before running any hardware comparison. If the curve has a long idle tail rather than a steady plateau, the cost case for self-hosting is considerably weaker than a simple per-token comparison would suggest, whatever the headline GPU price looks like on paper.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.