Services Case Studies Insights About Start a project →

Why every production AI system needs a model gateway.

AI Published August 18, 2026 9 min read

Why this matters right now

Stripe agreed to acquire OpenRouter, the startup that routes API calls across more than 400 AI models for roughly eight million developers, for more than $7 billion. That number is not the interesting part. What is interesting is what it says about where model routing sits in the stack now. Two years ago, a model gateway was a weekend project: a thin proxy that let a team swap between two or three providers without rewriting call sites. Today it is a category a payments company will spend billions of dollars to own, because whoever sits between an application and the model market controls the metering, the routing, and increasingly the billing for one of the fastest-growing line items on a technology budget.

For teams already running AI features in production, this is the moment the question stops being theoretical. If your application calls a single provider's SDK directly from a dozen places in the codebase, you already have the problem a gateway solves, you just have not paid the bill for it yet. That bill comes due the first time a provider has an outage during a product launch, the first time a finance leader asks why the AI line item jumped 40% in a month with no matching usage increase, or the first time a better, cheaper model ships and migrating to it means touching every call site in the application. A model gateway is the architectural answer to all three problems at once, and the acquisition is a signal that the market has decided this is infrastructure, not a nice-to-have.

What a model gateway actually does

Strip away the marketing and a model gateway is a reverse proxy with opinions. It sits between your application and every model provider you use, and it does the translation, routing, and bookkeeping work that would otherwise be scattered across your codebase. Instead of your application code knowing that provider A wants a slightly different message format than provider B, or that provider C throttles you at a certain rate, the gateway absorbs that complexity and exposes one consistent interface to everything upstream of it.

What this looks like in practice

  • One request schema, many providers underneath. Your application sends a single, consistent request shape, and the gateway translates it into whatever format the target provider expects, so a provider swap is a configuration change, not a code change.
  • Routing rules decide where a request actually goes. Rules can route by cost, by latency, by which provider is currently healthy, or by which model has the right capability for the task, all without the calling code knowing or caring.
  • Fallback chains absorb provider outages. When a provider is down or rate-limited, the gateway retries against a secondary provider automatically, turning what would be a customer-facing incident into a latency blip nobody notices.
  • Usage metering happens in one place. Every call passes through the same choke point, which is exactly why a payments company wanted this: the gateway is the natural place to measure, attribute, and eventually bill for AI consumption.

The build vs buy decision

Every team we talk to about this ends up choosing between three real options, and the right answer depends on how much of your traffic is AI-driven and how sensitive your data is. The first option is a hosted gateway such as OpenRouter, Portkey, or Martian, where you get provider breadth and routing intelligence on day one in exchange for every prompt and response passing through a third party's infrastructure. The second is a self-hosted open source gateway such as LiteLLM or Kong's AI Gateway, which gives you the same routing and metering patterns while keeping traffic inside your own network boundary, at the cost of running and patching another piece of infrastructure yourself. The third is the native routing layer offered by a cloud provider, such as cross-region inference on Bedrock or the gateway features inside Azure AI Foundry, which trades flexibility for tight integration with infrastructure you already operate.

The Stripe acquisition sharpens this decision rather than settling it. A hosted gateway that also happens to be owned by a billing and payments company means your model traffic and your AI spend data now live inside the same corporate entity, which is either a convenience or a concentration risk depending on your industry and your contracts. Teams handling regulated data, or anyone who has ever had to answer a data residency question from a customer's security team, should read that fact carefully before routing production traffic through any hosted gateway, regardless of who owns it.

  • Choose hosted if speed matters more than control. A startup validating whether an AI feature has legs should not spend its first quarter building routing infrastructure it may throw away.
  • Choose self-hosted if you already know AI is core to the product. Once AI spend is a material line item, the ongoing operational cost of running your own gateway is usually smaller than the compounding cost of vendor lock-in and opaque billing.
  • Choose a cloud-native option if you are already committed to one cloud. The integration benefits often outweigh the narrower model catalogue, especially for regulated workloads that already live entirely inside that cloud's compliance boundary.

Architecture patterns that hold up in production

Whichever option you choose, the patterns that actually earn their keep in production are consistent, and they map closely to lessons we have seen play out in cost and reliability work across client systems.

  • Route by task difficulty, not by habit. Most requests hitting a production AI system are routine. Sending every one of them to the most capable, most expensive model available is the single most common source of runaway AI spend, and a gateway is where that routing decision belongs, not scattered through application code.
  • Treat fallback chains as a reliability requirement, not an optimisation. A provider outage should degrade your service, not take it down. Define an explicit fallback order per use case ahead of time, and test it the same way you would test a database failover.
  • Cache at the gateway, not just in the application. Centralising caching at the routing layer means every consumer of the gateway benefits from a cache hit, instead of each team reimplementing its own caching logic inconsistently.
  • Tag every request with a cost centre. Attribute spend to the feature, team, or customer that generated it from day one. Retrofitting cost attribution after a budget surprise is far more painful than building it in from the start.
  • Instrument latency, cost, and error rate per provider and per model. A gateway that cannot tell you which upstream is degrading right now is not giving you the operational visibility the architecture is supposed to buy you.

Governance and the new billing entanglement

The part of this deal that deserves more scrutiny than it has gotten is what happens when the company routing your model traffic is the same company processing your payments. A gateway sees the content of every prompt and response that passes through it, which makes it one of the most sensitive aggregation points in a modern AI stack, arguably more sensitive than the database, because it sees the raw inputs and outputs before any application-level redaction happens. Folding that into a billing and payments platform raises legitimate questions about data handling boundaries, about whether routing decisions could ever be influenced by commercial arrangements with specific model providers, and about long-term pricing once a category consolidates around fewer, larger vendors.

None of that means avoid hosted gateways outright. It means the same governance rigour that applies to a payment processor or a cloud vendor now needs to apply to whichever company sits in front of your model calls. Read the data processing agreement. Confirm where prompts and responses are logged and for how long. Build your integration against an abstraction you control, so that if pricing or terms change after a future acquisition, migrating to a different gateway is a configuration change rather than a rewrite. The teams that will handle the next few years of consolidation in this space well are the ones who treated the gateway as infrastructure worth governing from the start, not as a convenience they bolted on.

Practical takeaways

The architecture question a model gateway answers was always going to matter once AI spend became a real line item, and a $7 billion acquisition just moved the timeline up for a lot of teams that were planning to deal with it later. Start by auditing how many places in your codebase call a model provider directly today, because that number is a reasonable proxy for how much pain a provider outage or a pricing change would cause you. If AI is a small, experimental part of your product, a hosted gateway will get you most of the benefit with the least engineering effort. If it is core to what you sell, the operational discipline of running your own routing layer, with cost attribution and fallback chains built in from day one, pays for itself the first time a provider has a bad week.

If you are weighing that build versus buy decision against a real budget and a real deadline, that assessment, including where a gateway fits your specific provider mix and compliance requirements, is exactly the kind of scoping work we do as part of our AI infrastructure consultancy. Start with an honest inventory of where your model calls live today, and let that inventory decide how much gateway you actually need.

Keep reading

Weighing a model gateway for your AI stack?

Start a conversation →
KT Solutions Assistant

Before we start, please share a few details so we can follow up with you.

Please enter your name and a valid email address.

End this conversation? Your chat will be emailed to us.