BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-09-11 · · Author:BazaarLink · LiteLLM · LiteLLM tutorial · Self-host LiteLLM · AI gateway · AI Gateway · LLM proxy · BYOK · Bring your own key · API key management

LiteLLM Tutorial and Self-Host Costs: Is BYOK an Option?

LiteLLM setup with Docker and virtual keys, plus the ongoing cost of self-hosting. If you only need one entry point for your own keys, BYOK needs no server.

Query date: 2026-09-14 (Taipei time). The LiteLLM commands and configuration formats are taken from its official documentation as of that date. The open-source project changes frequently, so check the official docs again before you start; BazaarLink's product behavior follows the actual configuration on our own site.

Most people who find this article are in one of two situations: they haven't set anything up yet and want to know what LiteLLM actually is and whether it's worth running; or they've already gone back and forth between docker compose and config.yaml for several rounds and are starting to wonder whether this proxy really has to be run at all.

This article serves both. The first half is a LiteLLM installation tutorial you can follow step by step. The second half works out the real math of self-hosting an AI gateway, including the cells that our own offering cannot cover.

What is LiteLLM? The short version

LiteLLM is an open-source AI Gateway / LLM proxy. It does not sell tokens. What it does is this:

You put your own OpenAI, Anthropic, Google and other API keys into its configuration file, and it exposes one OpenAI-compatible endpoint to the outside. Your code only knows this one URL. Changing models or providers is done in the configuration file, and the code stays the same.

Along the way it also does these things:

  • Virtual Keys: Instead of handing the real upstream key to every engineer or every agent, it issues a "virtual" key that can have a budget and rate limit, and can be revoked at any time
  • Load balancing and routing: Several upstreams can sit behind the same model name, and it splits traffic among them for you
  • Fallbacks: If one upstream goes down, it automatically switches to another
  • Usage and cost logs: Which key spent how much, stored in a database

Together, these four things are the core capabilities of an "AI gateway." The only question left is this: do you run these four things yourself, or let someone else run them? First, here is how to run it yourself. Then we'll come back to that question.

LiteLLM tutorial: up and running with Docker in ten minutes

Step 1: Write a minimal config.yaml

The core of LiteLLM is model_list. Each entry specifies "what it is called externally" and "which provider it actually calls and which key it uses":

model_list:
  - model_name: gpt-5.6-terra
    litellm_params:
      model: openai/gpt-5.6-terra
      api_key: os.environ/OPENAI_API_KEY

api_key: os.environ/OPENAI_API_KEY means the key is read from an environment variable. Do not write the key in plain text in the yaml, because this file can easily end up in version control by accident.

Step 2: Start the container

The simplest way to run it, without a database:

docker run \
  -v $(pwd)/litellm_config.yaml:/app/config.yaml \
  -e OPENAI_API_KEY=<your OpenAI API key> \
  -e LITELLM_MASTER_KEY=sk-1234 \
  -p 4000:4000 \
  docker.litellm.ai/berriai/litellm:latest \
  --config /app/config.yaml

Once it's running, the proxy is at http://localhost:4000. LITELLM_MASTER_KEY is the admin master key, and you must replace the example value sk-1234 before going live.

Step 3: Point your program at it

from openai import OpenAI

client = OpenAI(base_url="http://localhost:4000", api_key="sk-1234")
resp = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[{"role": "user", "content": "hello"}],
)

The model value is the model_name you defined in model_list, not the upstream's original name. At this point, it works.

Step 4 (what most people will need): Connect Postgres and enable virtual keys

The setup above has a limitation that many tutorials don't spell out: without a database, there are no virtual keys, no budget control, and no usage logs. The official documentation is direct about it: without a database, requests that carry a virtual key fail outright, and spending limits can only be set in each upstream provider's own dashboard.

So in practice you use the official compose file, which brings up the proxy and Postgres together:

curl -sSL https://docs.litellm.ai/docker-compose.yml | docker compose -f - up -d

After it starts, open http://localhost:4000/ui, create keys under "Virtual Keys," and set budgets and rates. Once this step is done, you actually have the four capabilities described earlier.

Step 5: Two things to do before going live

  1. Set LITELLM_SALT_KEY. This is the salt used to encrypt the upstream keys you store. The official docs list it as a required item before production. If it is not set, the protection your upstream keys get in the database is weaker than you think.
  2. Replace LITELLM_MASTER_KEY with a long random value. Whoever has the master key can create any virtual key.

Once these two keys are set, you have formally taken on the responsibility of "holding other people's keys." We'll come back to the weight of that later.

Advanced configuration: how to write load balancing and fallbacks

This is the section where people searching for "litellm tutorial" usually get stuck. The syntax below is from the official documentation (version of 2026-09-14).

Multiple upstreams under one name = load balancing

If you give two model_list entries the same model_name, LiteLLM will distribute traffic between them. You can set rpm / tpm on each to weight them:

model_list:
  - model_name: my-flash
    litellm_params:
      model: openai/gpt-5.6-luna
      api_key: os.environ/OPENAI_API_KEY
      rpm: 60
  - model_name: my-flash
    litellm_params:
      model: gemini/gemini-3.8-flash
      api_key: os.environ/GEMINI_API_KEY
      rpm: 60

The distribution strategy is set with routing_strategy under router_settings. The options are simple-shuffle (default), least-busy, usage-based-routing, and latency-based-routing.

Fallbacks

litellm_settings:
  fallbacks: [{"my-flash": ["gpt-5.6-terra"]}]
  context_window_fallbacks: [{"my-flash": ["gpt-5.6-terra"]}]

fallbacks defines where to fall back for general errors. context_window_fallbacks is a fallback specifically for "context too long" errors. These two must be configured separately. Many people only set the first one, and then still get errors when processing long documents.

At this point, a working LiteLLM setup is complete.

The five errors you're most likely to hit after setup

This section is for people who have it running but something still isn't working. Each of these is a configuration problem, not a bug in LiteLLM.

What you seeMost likely causeHow to fix it
Calls return 401The api_key your program sends is neither LITELLM_MASTER_KEY nor a valid virtual keyGet it working with the master key first, then switch to a virtual key; virtual keys require a database
model not foundThe model in your program is the upstream's original name (for example openai/gpt-5.6-terra), but LiteLLM only recognizes the model_name in model_listAlways use the model_name you defined
Virtual keys always fail, master key worksNo Postgres connectedStart the proxy plus database with the official compose file
Long documents return a context-too-long error with no fallbackOnly fallbacks is set, not context_window_fallbacksSet both; the latter handles context errors specifically
After a restart, upstream keys can't be read or decryptedLITELLM_SALT_KEY was changed, or it differs between deploymentsOnce set, the salt must stay fixed and be backed up carefully; changing it invalidates all stored keys

If all five of these check out and it still doesn't work, the next step is to check the proxy container's logs. Don't suspect the upstream first. In nine cases out of ten the problem lies between you and LiteLLM, not between LiteLLM and the upstream.

What follows is the part that no tutorial lists as a step.

When you should keep self-hosting

Let's start with the other side, because it matters more: in the three situations below, no hosted service should talk you out of self-hosting, including us.

  1. Request content cannot leave your data center. This is the most fundamental and irreplaceable reason to self-host. For finance, healthcare, government, or teams whose internal policies prohibit forwarding through third parties, self-hosting is the only option. A hosted gateway, no matter how it encrypts or promises not to retain data, still means your requests pass through someone else's machines.
  2. You need to change how it behaves. Custom routing logic, inserting your own middleware, modifying billing formulas, or connecting to internal model servers. The value of open source lies precisely here.
  3. You already have DevOps capacity, and this proxy is part of an existing system. If you already run a service mesh and an API gateway, the marginal cost of one more container is close to zero.

If you fall into one of these three categories, you can stop reading here. Continuing to self-host is the right decision, and the tutorial above is written for you.

The real cost of self-hosting (not just the "getting it running" part)

Getting LiteLLM running is indeed not hard; it's the ten minutes above. The real cost comes after that, and it doesn't appear in any comparison table.

What you take onWhy it isn't free
A process on the critical pathIt sits on the path of every request. If it goes down, all your AI features go down with it. A self-hosted proxy has no failover, and if it crashes at midnight, you are the one who gets up to fix it.
Version upgradesThe model catalog keeps changing. A proxy pinned to one version will quietly fail to recognize new models; keeping up means re-verifying everything with every upgrade. Someone has to decide whether to upgrade or not.
Other people's keys now live with youYou are now the custodian of every upstream provider key. Is LITELLM_SALT_KEY set? Is the database backup encrypted? Who can read it? From now on, all of these questions are yours.
ObservabilityUsage logs, cost attribution, and alerts: LiteLLM gives you the data, but you build the reports and alerts yourself.
Retry and rate-limit behaviorEvery upstream provider has different error codes. Wrong retry logic burns real money; no retries means users see errors.
Long streams during deploymentA stream that runs for 20 minutes gets cut if you do a rolling update of the proxy mid-stream. You need to plan a graceful drain, or every deployment is a gamble.
Key rotationEvery time you replace an upstream key, you must edit the configuration, restart, and verify again.

Each of these is not hard on its own. Together, they make up a standing operations job, and that job never ends.

Another path: bring your keys, without running the server yourself

If your only reason to self-host is this: "I already have my own API keys and want one unified entry point to manage them," then what you need is the "bring your own key" capability, not "running a server yourself." In LiteLLM these two are tied together, but they could be separated.

BYOK (Bring Your Own Key) provides that same LiteLLM layer. The difference is that the server is not yours.

You connect your own upstream accounts. From then on, requests to that provider are sent using your own key and billed to your own upstream account, and the platform does not charge for those requests. Setup takes three steps: enter a name and base URL, the system immediately makes one verification call and checks the returned model list, then you're done.

Mapping each LiteLLM feature

LiteLLM capabilityWhat it corresponds to in BYOK
model_list unified endpointYour own providers and platform-proxied models mixed, under the same base URL
Load balancing (multiple upstreams under one name)One model connected to more than one upstream route, with primary and fallback routes tracked separately for health
fallbacksTwo modes: seamless (if your key fails, that request is automatically completed through platform routing and billed at platform rates) or strict (fallback turned off, only your account is used)
Virtual keys + rate limitsEach API key can have a per-minute request cap
Usage logsBuilt in; shows which requests used your key, whether any retries happened, and what the failure codes were
Content filtering (you have to wire it up yourself with LiteLLM)Built in; BYOK requests also go through it, so violating requests are blocked before they ever reach your own upstream account

The last row deserves a bit more explanation. Suppose you connect your own OpenAI key and someone on your team sends a violating request. That request would be sent to the upstream using your account. With self-hosted LiteLLM, you have to wire up moderation yourself; with BYOK, the platform's filtering layer blocks it first.

How to call it

curl https://api.bazaarlink.ai/v1/chat/completions \
  -H "Authorization: Bearer $BAZAARLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "messages": [{"role": "user", "content": "hello"}]
  }'

It is OpenAI-compatible, so existing SDKs only need the base URL changed. This is the same action as pointing your program at LiteLLM above; the only difference is the address.

An honest comparison table

The point of this table is the "cannot do" entries on the right. Without those cells, the comparison below would have no value.

Self-hosted LiteLLMBYOK (hosted gateway)
Use your own upstream keys✅✅
OpenAI-compatible unified endpoint✅✅
Mix multiple upstream providers✅✅
Failover strategyWrite it yourself in yamlChoose one of two modes
Usage visibilityConnect Postgres and build your own reportsBuilt in
Per-minute request cap✅ (requires a database)✅
Content filteringWire it up yourselfBuilt in; BYOK also goes through it
Requests never leave your data center✅❌ Cannot do
Modify source code / custom routing logic✅❌ Cannot do
Control spending on your own keys with a dollar cap✅ (virtual key budgets)❌ Cannot do (see below)
Operations responsibilityYouUs

⚠️ One limitation you should know about first

For requests using your own keys, we record a cost of 0, because that money goes directly from you to your own upstream provider, and we don't handle it.

The practical consequence: you cannot use a dollar cap like "$100 per month" to control usage on your own keys, because those requests don't accumulate into any dollar amount. Spending caps you have already set will still block requests (BYOK does not bypass that gate), but they will not increase because you used your own keys.

To limit usage on your own keys, what you can currently use is the per-minute request cap, plus setting quotas for that account in your own upstream dashboard. If "setting a dollar budget per virtual key" is the main reason you self-host, then BYOK cannot yet replace LiteLLM's virtual keys, and we won't gloss over that.

Also, usage on your own keys is not counted toward the platform's usage rebate, which applies to requests billed through the platform.

When this path makes more sense

Combining the two tables above, the decision is simple:

  • If your reason to self-host is data residency or the need to change code → keep self-hosting, and use the tutorial above
  • If your reason to self-host is using your own keys plus one unified entry point → you don't need to run a server for this
  • If your reason to self-host is capping each person's budget by dollar amount → self-hosting is still the better fit for now
  • If your reason to self-host is "everyone else is self-hosting" → that is not a reason. Every extra component is extra operations responsibility, and that responsibility doesn't appear in any comparison table. It shows up on some weekend instead

How to tell whether your LiteLLM has become a burden

Self-hosting is not a mistake. Failing to notice that it has started to consume you is the problem. Answer the five questions below. If three or more of them give you pause, it's worth reconsidering:

  1. In the past three months, has any AI feature outage been caused by the proxy itself (crashes, OOM, behavior changed after an upgrade) rather than by the upstream?
  2. When a new model comes out, how many days pass between "available upstream" and "usable by your team"? Who are you waiting on during those days?
  3. Besides one person, does anyone on your team know where LITELLM_SALT_KEY is stored and how to rotate it?
  4. When was the last time someone seriously looked at the proxy's usage reports? Or is the database just growing with nobody looking at it?
  5. If the person who set up the proxy left tomorrow, would anyone dare touch it?

None of these five questions asks whether LiteLLM is easy to use. They ask: is this proxy now a part of your system, or a single point of failure within it?

Practical steps to move from LiteLLM to BYOK

If the five questions above have led you to decide to stop running it yourself, here is a migration path that requires no downtime. The core principle: run in parallel first, then switch, and only tear down last.

Step 1: List out your model_list

Open your config.yaml and list which upstream provider and which key each model_list entry uses. This list is the set of providers you need to bring into BYOK. It is usually shorter than you expect; most teams actually use only two or three providers.

Step 2: Connect each provider to BYOK

After logging in, go to /keys/byok and fill in each provider one by one: name, that provider's base URL, and your own API key. The system immediately makes one verification call and checks the returned model list. It counts as connected only when it passes, not just when the form is filled in.

Step 3: Decide the fallback mode

The fallbacks in LiteLLM map to a single switch here. If your original fallback routed to another one of your own keys, choose strict mode and connect both providers. If your original setup was "use my key as much as possible, and if it truly fails, that's acceptable," choose seamless: when your key fails, that request is completed through platform routing and billed at platform rates.

Step 4: Change only one environment variable

Change OPENAI_BASE_URL in your program from http://localhost:4000 to https://api.bazaarlink.ai/v1, and replace OPENAI_API_KEY with your BazaarLink key. Just check the model names against the mapping. This step is the same action as when you originally pointed your program at LiteLLM.

Step 5: Run in parallel for a week and compare the usage logs on both sides

Don't tear down LiteLLM yet. Send part of the traffic through the new entry point and compare the usage logs on both sides: which requests used your own key, whether any fallbacks occurred, and what the failure codes were. Once a week passes with no surprises, move the rest of the traffic over, and only then stop the container.

After the move, every row in the "real cost of self-hosting" table from the first section no longer belongs to you. The only thing that stays with you is your own upstream keys and their bills, which were always yours.

Try it first

Connecting your first key of your own takes about five minutes; after that, requests to that provider go through your own account. The entry point is Bring Your Own Key (BYOK). After logging in, configure it at /keys/byok.

If you don't have an account yet and just want to confirm compatibility first, there is a free tier you can try: Qwen3.7 Flash is free, 10 requests per minute, 50 per day, no credit card required. That's enough to change the base URL and confirm that your SDK works.

Further reading: What's the difference between an AI API gateway, a router, and a relay station, A unified AI model entry point for engineering teams, How to tell whether an AI API relay station is trustworthy.

FAQ

What is LiteLLM?

LiteLLM is an open-source AI gateway (LLM proxy). It does not sell tokens. Instead, it centrally manages your own API keys from OpenAI, Anthropic, Google and other providers, and exposes a single OpenAI-compatible endpoint. It also provides virtual keys (with budgets and rate limits), load balancing and failover across multiple upstreams, and usage logs. Your code only needs one URL, and switching models is done in the config file.

How do I install LiteLLM? What is the shortest path?

Write a config.yaml (model_list lists the public name, the actual model, and the key read from an environment variable), run docker run with the config file mounted and LITELLM_MASTER_KEY set, expose port 4000, and point base_url to http://localhost:4000 in your code. To use virtual keys, budgets and usage logs, you must connect Postgres; the official docker-compose file brings everything up at once. Before going live, be sure to set LITELLM_SALT_KEY and replace the master key.

Can I use LiteLLM without a database?

Yes, but you only get the unified endpoint. Without a database there are no virtual keys (requests using virtual keys will fail), no budget control, and no usage logs. Spending limits can only be set in each upstream provider's own dashboard. Most teams actually need to connect Postgres.

What is the biggest difference between BYOK and self-hosted LiteLLM?

They do the same layer of work: your own upstream keys, one unified entry point, failover, and usage logs. The difference is that the server is not yours. You don't operate the proxy, don't manage key encryption, and don't build your own reporting. But BYOK cannot do three things: keep requests inside your own data center, modify source code to customize routing, and control spending on your own keys with a dollar cap.

If I use BYOK, how much will you charge me?

Requests using your own keys are not charged a platform fee; that money goes directly from you to your own upstream provider. Only if your key fails and you have selected "seamless" mode does that request get rerouted through platform routing and completed under normal platform billing. In "strict" mode there is no fallback; a failure is a failure.

What happens if my own key fails?

It depends on the mode you choose. Seamless (default): the request is automatically rerouted through platform routing and billed at platform rates, so the request does not get stuck. Strict: automatic fallback is turned off, the request only uses your own account, and an error is returned on failure. If you want to be sure every request uses your own key, choose strict.

Can I set a dollar budget per person, like LiteLLM's virtual keys?

Not for BYOK usage at the moment. Requests using your own keys are recorded with a cost of 0 (you pay the upstream directly), so they do not add to dollar caps. Spending caps you have already set still block requests, but they do not increase because of BYOK usage. What you can use is the per-minute request limit and quotas set in your own upstream dashboard. If a dollar budget is the main reason you self-host, self-hosting is still the better fit for now.

Can I mix my own keys with models provided by the platform?

Yes. Providers you bring and models proxied by the platform share the same base URL and the same BazaarLink API key, so from the code's point of view it is one endpoint. In the usage logs you can see which requests went through your own key.

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
Claude Code · Claude Code update · changelog · mods · permission rules
Claude Code 2.1.289: 27 Changes, Permission-Rule Fixes and Mods Stability
Claude Code · Claude Code mods · plugins · TypeScript · AI coding tools
Claude Code Mods: Customize Behavior and UI with TypeScript
web search API · AI agent · LLM · Hermes Agent
Hermes Agent Web Search: Use a BazaarLink Model with :online
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.