LLM Gateways Reviewed: Adopt One at Your Second Provider
Routing, failover, caching, and the spend attribution nobody else provides — against a new single point of failure.
Jump to 5 sections
Quick answer: An LLM gateway sits between your application and one or more model providers, handling routing, failover, caching, rate limiting, and spend tracking behind a single API. It is worth adopting once you use more than one provider or need centralized cost control. For a single-provider application it adds a hop and a dependency for little return.
Gateways became mainstream infrastructure during 2026 for an unglamorous reason: teams ended up with three providers, four models, and no single place to see what any of it cost.
This review covers the LLM gateway category — what these systems do, which features earn their place, the latency and reliability cost of adding a hop, and when a thin internal wrapper beats a product.
What an LLM gateway actually does
It presents one API surface and fans out to several behind it. Your application calls the gateway; the gateway decides which provider and model handles the request, retries or fails over on error, and records what it cost.
The value is centralization rather than any individual capability. Without a gateway, routing logic, retry policy, key management, and cost attribution get reimplemented in every service that calls a model — inconsistently, and usually without the cost part.
We have reviewed enough of these internal reimplementations at Model Drop to see the pattern clearly: the third service to call a model copies the retry logic from the second, which copied it from the first, and by the fourth service the retry behavior, the timeout values, and the error handling have all quietly diverged. Nobody decided that on purpose. A gateway is partly a technical fix and partly a way of stopping that drift before it compounds.
Four functions form the core. Request routing by model, cost, or policy. Failover when a provider returns errors or times out. Caching of repeated requests. And centralized spend tracking with per-team or per-feature attribution.
That last one is frequently the actual reason a gateway gets adopted. When the monthly inference bill becomes large enough to attract attention, somebody asks which team spent it, and a gateway is the only place that answer lives — the magnitudes involved are in our breakdown of LLM API pricing across the 2026 tiers.
Which features earn their place
Not all of them, and the marketing does not distinguish.
| Feature | Verdict | Why |
|---|---|---|
| Unified API | Earns it | Removes per-provider SDK sprawl |
| Failover routing | Earns it | Provider outages are real |
| Spend attribution | Earns it | Nothing else provides it |
| Exact-match caching | Earns it | Free wins on repeated requests |
| Semantic caching | Be careful | Similar is not equivalent |
| Automatic model selection | Be careful | Opaque quality changes |
Semantic caching deserves the scrutiny. Returning a cached response for a semantically similar but not identical query sounds efficient and means your users sometimes receive an answer to a question they did not ask. The similarity threshold is a quality-versus-savings dial, and it should be set deliberately rather than left at a default.
Automatic model selection has a related problem. A gateway that silently routes to a cheaper model when it judges a request simple is making a quality decision on your behalf, with no signal in your logs unless you capture which model served each request. If you use this, log the routing decision.
Provider-side prompt caching is a separate mechanism worth not confusing with gateway caching. The provider discount applies to a stable prompt prefix and is billed at a reduced rate on cache reads; a gateway cache avoids the call entirely. They compose, but a gateway that reorders or rewrites prompts can silently break the prefix stability that the provider discount depends on — a failure mode that shows up as a bill increase with no behavior change.
Key management belongs on the feature list too, and it is the quiet operational win. Holding provider credentials in one place, rotating them without redeploying every service, and scoping them per team is genuinely useful, and it is the kind of thing standards frameworks like the NIST AI Risk Management Framework treat as basic governance rather than an enhancement.
Exact-match caching, by contrast, is uncomplicated and free. Identical prompts return identical cached responses, which is straightforwardly correct for deterministic configurations.
Pricing across hosted gateways is usually a small percentage markup on top of provider cost, commonly in the 3-8% range, plus sometimes a flat per-seat or per-request fee once volume passes a threshold. On a team spending $8,000 a month across providers, an 5% gateway markup adds roughly $400 — a cost that's easy to justify once you can finally answer which feature is driving that spend, and hard to justify if nobody asks the question at all.
What adding a hop costs you
Latency and a dependency, and the second is the one to think about.
A well-implemented gateway adds single-digit milliseconds, which is negligible against model response times measured in hundreds of milliseconds or seconds. A poorly-placed one — in a different region from both your application and the provider — can add considerably more, sometimes 50-100ms round trip once cross-region routing is involved. Measure it rather than assuming; a gateway deployed in the wrong AWS region can turn an invisible overhead into a noticeable one on every single request.
The dependency is structural. Every model call now traverses a component that can fail, and a gateway outage takes down inference entirely, including the failover logic that was supposed to protect you. That is an uncomfortable topology, and it argues for self-hosting the gateway or choosing one with a documented uptime record and a direct-to-provider bypass path.
Prompt transit is the third consideration. A hosted gateway sees every prompt and response, which is the same data-handling question that arises with observability tooling — covered in our review of LLM observability tools. Self-hosted gateways avoid it entirely, and several mature options can be run in your own infrastructure.
For a team in a regulated industry, prompt transit is frequently the deciding factor rather than a footnote. Healthcare and financial services teams we have talked with generally rule out a third-party hosted gateway outright the moment patient or account data appears in a prompt, and go straight to a self-hosted option even though it costs more in operational overhead. That is a defensible default, not an overcautious one.
When should you actually adopt an LLM gateway?
Adopt a gateway once you use two or more model providers, or once several services call models with no shared way to track spend. Below that — one provider, one model, one calling service — a gateway adds a hop and a dependency for problems you do not have yet. Write a few lines of retry logic and move on.
We have watched this threshold get crossed by accident more often than by decision: a team adds a second provider as a stopgap during an outage, never removes it, and six months later nobody can say which of the two is serving which fraction of traffic. That drift is usually the actual trigger, not a deliberate architecture review.
Two or more providers, or several services calling models independently, and the calculation inverts. At that point routing logic and cost attribution are being reimplemented repeatedly, and centralizing them pays for itself quickly. Multi-provider routing is now standard practice for availability as much as cost, as we discuss in our comparison of inference providers.
Open standards reduce the lock-in question considerably. Most gateways now expose an interface compatible with the widely-adopted chat completions shape, and several emit telemetry using the OpenTelemetry semantic conventions for generative AI. That means migrating between gateways, or removing one, is usually a configuration change rather than a rewrite — which lowers the cost of getting this decision wrong.
Between those, a thin internal wrapper is legitimate and underrated. A shared library that normalizes provider differences, applies retry policy, and emits cost metrics gives you most of the benefit with no new runtime dependency. At Model Drop we would reach for that before a product in any single-provider deployment.
Building that wrapper usually takes a small team a few days, not weeks — the hard part is agreeing on the interface once, not the implementation. Teams that skip this step and let every service call providers directly almost always end up building the wrapper anyway, just later and under more pressure, once the third or fourth service exposes how inconsistent the earlier ad hoc calls had become.
The bottom line on LLM gateways
Gateways centralize routing, failover, caching, and spend attribution, and the last of those is usually what triggers adoption. The threshold is a second provider. Exact-match caching and failover earn their place; semantic caching and automatic model selection are quality decisions that need explicit configuration and logging.
Your next step: check whether you can currently answer which team or feature spent last month's inference budget. If not, that gap is what a gateway fixes, and it is the strongest argument for one.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.