Platforms

LLM Gateways Reviewed: Adopt One at Your Second Provider

Routing, failover, caching, and the spend attribution nobody else provides — against a new single point of failure.

Dana Kwon

Contributing Reviewer

Published 5 min read
Close-up image of ethernet cables plugged into a network switch, showcasing IT infrastructure.
Jump to 5 sections

Quick answer: An LLM gateway sits between your application and one or more model providers, handling routing, failover, caching, rate limiting, and spend tracking behind a single API. It is worth adopting once you use more than one provider or need centralized cost control. For a single-provider application it adds a hop and a dependency for little return.

Gateways became mainstream infrastructure during 2026 for an unglamorous reason: teams ended up with three providers, four models, and no single place to see what any of it cost.

This review covers the LLM gateway category — what these systems do, which features earn their place, the latency and reliability cost of adding a hop, and when a thin internal wrapper beats a product.

What an LLM gateway actually does

It presents one API surface and fans out to several behind it. Your application calls the gateway; the gateway decides which provider and model handles the request, retries or fails over on error, and records what it cost.

A close-up of a red traffic light set against a clear blue sky, conveying safety and caution.

The value is centralization rather than any individual capability. Without a gateway, routing logic, retry policy, key management, and cost attribution get reimplemented in every service that calls a model — inconsistently, and usually without the cost part.

We have reviewed enough of these internal reimplementations at Model Drop to see the pattern clearly: the third service to call a model copies the retry logic from the second, which copied it from the first, and by the fourth service the retry behavior, the timeout values, and the error handling have all quietly diverged. Nobody decided that on purpose. A gateway is partly a technical fix and partly a way of stopping that drift before it compounds.

Four functions form the core. Request routing by model, cost, or policy. Failover when a provider returns errors or times out. Caching of repeated requests. And centralized spend tracking with per-team or per-feature attribution.

That last one is frequently the actual reason a gateway gets adopted. When the monthly inference bill becomes large enough to attract attention, somebody asks which team spent it, and a gateway is the only place that answer lives — the magnitudes involved are in our breakdown of LLM API pricing across the 2026 tiers.

Which features earn their place

Not all of them, and the marketing does not distinguish.

Detailed view of a car's speedometer and tachometer displaying speed and RPM.
Which features earn their place
FeatureVerdictWhy
Unified APIEarns itRemoves per-provider SDK sprawl
Failover routingEarns itProvider outages are real
Spend attributionEarns itNothing else provides it
Exact-match cachingEarns itFree wins on repeated requests
Semantic cachingBe carefulSimilar is not equivalent
Automatic model selectionBe carefulOpaque quality changes

Semantic caching deserves the scrutiny. Returning a cached response for a semantically similar but not identical query sounds efficient and means your users sometimes receive an answer to a question they did not ask. The similarity threshold is a quality-versus-savings dial, and it should be set deliberately rather than left at a default.

Automatic model selection has a related problem. A gateway that silently routes to a cheaper model when it judges a request simple is making a quality decision on your behalf, with no signal in your logs unless you capture which model served each request. If you use this, log the routing decision.

Provider-side prompt caching is a separate mechanism worth not confusing with gateway caching. The provider discount applies to a stable prompt prefix and is billed at a reduced rate on cache reads; a gateway cache avoids the call entirely. They compose, but a gateway that reorders or rewrites prompts can silently break the prefix stability that the provider discount depends on — a failure mode that shows up as a bill increase with no behavior change.

Key management belongs on the feature list too, and it is the quiet operational win. Holding provider credentials in one place, rotating them without redeploying every service, and scoping them per team is genuinely useful, and it is the kind of thing standards frameworks like the NIST AI Risk Management Framework treat as basic governance rather than an enhancement.

Exact-match caching, by contrast, is uncomplicated and free. Identical prompts return identical cached responses, which is straightforwardly correct for deterministic configurations.

Pricing across hosted gateways is usually a small percentage markup on top of provider cost, commonly in the 3-8% range, plus sometimes a flat per-seat or per-request fee once volume passes a threshold. On a team spending $8,000 a month across providers, an 5% gateway markup adds roughly $400 — a cost that's easy to justify once you can finally answer which feature is driving that spend, and hard to justify if nobody asks the question at all.

What adding a hop costs you

Latency and a dependency, and the second is the one to think about.

Close-up of professionals reviewing financial graphs at a business meeting.

A well-implemented gateway adds single-digit milliseconds, which is negligible against model response times measured in hundreds of milliseconds or seconds. A poorly-placed one — in a different region from both your application and the provider — can add considerably more, sometimes 50-100ms round trip once cross-region routing is involved. Measure it rather than assuming; a gateway deployed in the wrong AWS region can turn an invisible overhead into a noticeable one on every single request.

The dependency is structural. Every model call now traverses a component that can fail, and a gateway outage takes down inference entirely, including the failover logic that was supposed to protect you. That is an uncomfortable topology, and it argues for self-hosting the gateway or choosing one with a documented uptime record and a direct-to-provider bypass path.

Prompt transit is the third consideration. A hosted gateway sees every prompt and response, which is the same data-handling question that arises with observability tooling — covered in our review of LLM observability tools. Self-hosted gateways avoid it entirely, and several mature options can be run in your own infrastructure.

For a team in a regulated industry, prompt transit is frequently the deciding factor rather than a footnote. Healthcare and financial services teams we have talked with generally rule out a third-party hosted gateway outright the moment patient or account data appears in a prompt, and go straight to a self-hosted option even though it costs more in operational overhead. That is a defensible default, not an overcautious one.

When should you actually adopt an LLM gateway?

Adopt a gateway once you use two or more model providers, or once several services call models with no shared way to track spend. Below that — one provider, one model, one calling service — a gateway adds a hop and a dependency for problems you do not have yet. Write a few lines of retry logic and move on.

We have watched this threshold get crossed by accident more often than by decision: a team adds a second provider as a stopgap during an outage, never removes it, and six months later nobody can say which of the two is serving which fraction of traffic. That drift is usually the actual trigger, not a deliberate architecture review.

Detailed view of Ethernet and VGA ports on a server highlighting connectivity features.

Two or more providers, or several services calling models independently, and the calculation inverts. At that point routing logic and cost attribution are being reimplemented repeatedly, and centralizing them pays for itself quickly. Multi-provider routing is now standard practice for availability as much as cost, as we discuss in our comparison of inference providers.

Open standards reduce the lock-in question considerably. Most gateways now expose an interface compatible with the widely-adopted chat completions shape, and several emit telemetry using the OpenTelemetry semantic conventions for generative AI. That means migrating between gateways, or removing one, is usually a configuration change rather than a rewrite — which lowers the cost of getting this decision wrong.

Between those, a thin internal wrapper is legitimate and underrated. A shared library that normalizes provider differences, applies retry policy, and emits cost metrics gives you most of the benefit with no new runtime dependency. At Model Drop we would reach for that before a product in any single-provider deployment.

Building that wrapper usually takes a small team a few days, not weeks — the hard part is agreeing on the interface once, not the implementation. Teams that skip this step and let every service call providers directly almost always end up building the wrapper anyway, just later and under more pressure, once the third or fourth service exposes how inconsistent the earlier ad hoc calls had become.

The bottom line on LLM gateways

Gateways centralize routing, failover, caching, and spend attribution, and the last of those is usually what triggers adoption. The threshold is a second provider. Exact-match caching and failover earn their place; semantic caching and automatic model selection are quality decisions that need explicit configuration and logging.

Your next step: check whether you can currently answer which team or feature spent last month's inference budget. If not, that gap is what a gateway fixes, and it is the strongest argument for one.

Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.

What does an LLM gateway do?
It presents one API surface and fans out to several providers behind it, handling request routing by model or cost, failover on errors and timeouts, response caching, and centralized spend tracking with per-team attribution. The value is centralization rather than any single capability.
When should I adopt an LLM gateway?
At your second provider, or when several services call models independently. With one provider and one service, a gateway adds a hop and a dependency to solve problems you do not have — twenty lines of retry logic covers it. Past that threshold, centralizing pays for itself quickly.
Is semantic caching safe to enable?
Treat it carefully. Returning a cached response for a semantically similar but not identical query means users sometimes receive an answer to a question they did not ask. The similarity threshold is a quality-versus-savings dial that should be set deliberately rather than left at a default.
Does a gateway add meaningful latency?
A well-implemented one adds single-digit milliseconds, negligible against model responses measured in hundreds of milliseconds. A gateway placed in a different region from both your application and the provider can add considerably more, so measure it in your own topology rather than assuming.
What is the main risk of using a gateway?
It becomes a single point of failure. Every model call traverses a component that can fail, and a gateway outage takes down inference entirely, including the failover logic meant to protect you. Self-hosting the gateway or choosing one with a direct-to-provider bypass path mitigates this.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons