AI API Gateway vs Router vs Relay: Differences and Choices
Searching 'API Gateway' mostly returns Kong and APISIX. This guide separates AI API gateways, model routers and relays with six criteria, including BazaarLink.
Let me point out something that could waste an hour of your time: if you search "API Gateway" on Google, nearly all of the top ten results are enterprise API management platforms: Kong, APISIX, the API Gateway services of various cloud providers, and solution pages from system integrators. They manage the APIs you expose to others: traffic control, authentication and microservice routing.
That is not what you need. You need the layer that manages the APIs you call out to models. The two share the word "gateway" but are unrelated markets. This article covers only the latter.
Three terms, three different things
Within the scope of "managing the model APIs you call out to," the market has three things that are often lumped together:
AI API gateway (AI Gateway / LLM Gateway)
A service that sits between your program and the various model providers. Your program recognizes only one endpoint and one request format; the gateway is responsible for routing requests to the right provider, switching to another route when a provider has problems, issuing keys, recording usage, and blocking policy-violating content. Its value is that you do not have to maintain this layer yourself. You can self-host it (LiteLLM is the most common open-source choice) or use a hosted one.
Model router
Strictly speaking, a feature inside a gateway, but some products promote it as a standalone product: based on the content of each request, it automatically decides which model to send it to, sending simple questions to cheaper models and complex ones to more expensive models. The focus is the algorithm for choosing a model, not infrastructure such as keys, failover or billing.
If you already know which model you want to use, a router is worth close to nothing to you; it only makes sense when your traffic is mixed and you want to save money without classifying it yourself.
Relay (中轉站)
This term comes from the Chinese-speaking world. The essence of a relay is "reselling quota": the operator obtains access to various models in its own way and then sells it to you at a lower price. It usually also gives you an OpenAI-compatible endpoint, which looks very much like a gateway from the outside.
The difference lies in who you pay, and whether you can verify that what you receive is genuine. A gateway usually lets you see (and sometimes even use your own) upstream account; a relay's upstream is a black box, and what you buy is the claim that "this operator says it is model X."
This does not mean every relay has problems. But "unverifiable" is itself a cost, and below we explain how to verify it.
Six criteria: ask any provider these questions
Whether a provider calls itself a gateway, a router or a relay, asking these six questions will reveal its actual category. We have filled in our own answers as well, including the ones we cannot meet.
| Criterion | Why it matters | BazaarLink's answer |
|---|---|---|
| 1. Who gets paid? | Paying a model provider (through a gateway procurement) and paying a reseller carry different risk structures | Models procured by the platform: paid to us, with prices based on each provider's official list price; bring your own key (BYOK): paid to your own upstream, with no markup from us |
| 2. Can you use your own upstream account? | If yes, this layer is "routing" rather than "reselling" | Yes. After connecting your key, requests for that provider go through your own account |
| 3. What happens when an upstream goes down? | Is the request resent, or is the same request taken over and completed? Are you told that it happened? | One model can be served by more than one upstream route; when the primary route fails the health check, the same request is completed on a backup route. Retries happen before you receive any content; failures and retries remain in your usage records and are not silently swallowed |
| 4. What if the connection breaks after content has started? | This is the case most easily fooled by an answer that "looks complete" | It will not reconnect through another provider for you. You see a clear error, not an answer cut off midway; such requests are not charged. This describes the mechanism; it is not an availability guarantee |
| 5. How do I confirm I got the model I paid for? | Any endpoint can answer "I am GPT"; self-claims are not evidence | We do not act as the judge ourselves: the relay detection tool runs 84 standard checks against any endpoint, compares behavioral fingerprints, trusts zero self-claims, and writes "cannot determine" when the evidence is insufficient. You can use it to test us |
| 6. Who handles my requests? | Compliance red line | Our machines. If your policy is "no forwarding through third parties," we are not an option; you should self-host |
Row 6 is something we cannot do; row 4 is something we deliberately do not do. A comparison table with only checkmarks is usually made to sell something, not to help you choose.
Three situations Taiwanese teams actually encounter
Scenario 1: Two or three engineers, just moved up from calling the OpenAI SDK directly
Your pain points are one key per person, scattered monthly bills, and nobody knows which agent spent how much.
What you need is the gateway's key management and usage attribution, not a router. You can self-host LiteLLM to achieve this, but that proxy then sits on your critical path from then on. We lay out the real costs of self-hosting in this article. For small teams, nobody usually takes on that ongoing operations work.
Scenario 2: You already have an enterprise contract or existing quota to use up
You have quota under an OpenAI or Google contract, but also want to use models from other providers without maintaining three SDKs in your program.
This is the most typical use of bring your own key (BYOK): connect your contract account, so requests for that provider go through your own account with no markup; other models go through the platform. The program side has only one base URL.
Note one thing: for BYOK requests, the fee we record is 0 (you pay the upstream directly), so you cannot use the platform's spending cap to limit BYOK spending; what you can use are the per-minute request limit and the quota in your own upstream dashboard. We have written this restriction on the BYOK page and in the article, and we do not work around it.
Scenario 3: Compliance requires that requests never leave your data center
Self-host. No hosted gateway should talk you out of this, including us. LiteLLM is a reasonable starting point; the installation steps are here.
How to verify whether a gateway or relay is honest
The fifth criterion above says self-claims are not evidence. Here is how to get evidence. The principle comes down to one sentence: do not ask it who it is; make it do work, then compare how it does the work against known models.
Our relay detection works this way, in five stages, and each stage can conclude "insufficient evidence, cannot determine":
- Run the questions: up to 86 standard questions, first confirming that the questions were actually answered. Missing questions are not just a shortage of data; the remaining evidence may happen to lean one way
- Identify the provider: compare the style and habits of the answers, not what it claims to be. When it "claims one provider but behaves like another," treat the claim as a slip of the tongue and trust the behavior
- Identify the model: which specific model within that provider. So-called "dumbing down" mostly happens at this layer: the family stays the same, but it is swapped for a cheaper, smaller model from the same provider
- Twin re-verification: for highly similar close-relative models, a separate set of questions is sampled again, because the previous stage has almost no discriminating power between them
- Compare the answers: measured results vs the vendor's claims, with eight possible conclusions, including "ambiguous"
Use a revocable, low-quota test key and paste the base URL of any endpoint to run it. That includes ours. A gateway that does not dare let you test it already has an answer for the fifth criterion.
The short version
- Found Kong, APISIX, or a cloud API Gateway → That manages your own APIs; close the tab
- You need the layer that manages "the APIs you call out to models" → AI gateway
- A provider that leads with "automatically picks the model for you" → Router, which is only valuable if you are not sure which model to use
- A provider that is cheaper but you pay the operator and cannot see the upstream → Relay; run the detection tool once before deciding
- Your policy says requests may not leave the data center → Self-host; do not let anyone talk you out of it
If you want a gateway and do not want to run it yourself: connecting your first key takes about five minutes; see bring your own key (BYOK). To see how failover actually works first, read upstream failover.
FAQ
How is an AI API gateway different from an API Gateway?
An API Gateway (Kong, APISIX, and the API gateway services of various cloud providers) manages the APIs you expose to others: traffic control, authentication and microservice routing. An AI API gateway (AI Gateway / LLM Gateway) manages the APIs you call out to model providers: a unified endpoint, switching providers without code changes, failover, issuing keys and recording usage. The two share the word 'gateway' but are unrelated markets, so they are easy to confuse when searching.
What is the difference between an AI gateway, a model router and a relay?
A gateway is the infrastructure layer between your program and the various model providers, responsible for a unified endpoint, failover, keys and usage. A router is, strictly speaking, one feature of a gateway: it automatically chooses a model for each request. Some products promote it as a standalone product. A relay comes from the Chinese-speaking world; in essence it resells quota: the relay station gains access and sells it to you at a lower price, while the upstream is a black box, so what you buy is the claim that 'this is model X.' The biggest dividing lines among the three are who you pay, whether you can use your own upstream account, and whether you can verify that the model you receive is genuine.
How do you verify that the model a relay or gateway provides is genuine?
Do not ask it who it is; make it do work. Send a set of standard questions to the endpoint, compare the behavioral characteristics of the answers against baselines of known models, identify the family first and then the specific model, run a separate set of questions to re-verify highly similar close relatives, and finally compare the measured results against what the vendor claims. Self-claims are the easiest to fake and cannot serve as evidence. BazaarLink's relay detection tool works this way: you can run it against any endpoint, including ours, with a revocable low-quota key.
When should you not use a hosted AI gateway?
There are three cases: requests must not leave your own data center (a compliance red line, since a hosted gateway passes requests through someone else's machines no matter how they are encrypted); you need to change the gateway's behavior, such as customizing routing logic or inserting your own middleware; or you already have DevOps capacity and this proxy is part of an existing system. In these cases, self-host; LiteLLM is a common starting point.
What is the difference between BYOK and buying models through the gateway?
The party you pay is different. Requests using your own key go out through your own upstream account; the bill belongs to your upstream, and the platform adds no markup. Models the platform procures are paid to the platform, with prices based on each provider's official list price. The two can be mixed, using the same base URL. Note that the platform records a cost of 0 for BYOK requests, so the platform's spending cap cannot control BYOK spending; what you can use instead are the per-minute request limits and the quota in your upstream provider's dashboard.
When an upstream goes down, does the gateway resend the request or take it over?
This is a question you should ask any gateway. BazaarLink's approach: one model can be served by more than one upstream route; when the primary route fails the health check, the same request is completed on a backup route instead of starting over. Retries happen before you receive any content, and an answer that has already started streaming will not be switched to another provider midway. If the connection breaks after content has started, it does not reconnect through another provider; you see a clear error, and such requests are not charged. This describes the mechanism; it is not an availability guarantee.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API