SillyTavern Taiwan API: Free Tier, Models, Chat Costs
Set SillyTavern to a Taiwan API via Custom OpenAI-compatible: endpoint, key, model. Free-tier limits, fallback rules, cost example, and troubleshooting order.
Settings and price check date: 2026-09-17 (Taipei time). This article checks the official SillyTavern documentation and this site's live model catalog. It did not re-run an installation or a paid conversation test. Interface names depend on the version you have installed. Connect with a short message first, then import the full character card.
If this is your first time with "Tavern" (酒館), see Tavern AI Beginner's Guide first. This article assumes you can already open SillyTavern and need to handle API connection, model selection, and the cost of long conversations.
Connecting SillyTavern to a Taiwan API: Three Steps
SillyTavern is a chat frontend and does not come with any model account. You set up characters and conversations in the frontend, and the replies are generated by the model you connect. To connect to this site, use its custom OpenAI-compatible source:
Step 1: Choose the Connection Type
Open the API Connections icon (the plug), set the API type to Chat Completion, and set the source to Custom (OpenAI-compatible). This is the setup used in this article; it does not mean every other source option can never be configured for a custom service.
Step 2: Enter the Base URL and API Key
Custom Endpoint: https://api.bazaarlink.ai/v1
Custom API Key: the API key you created on this site
Create your key on the API Keys page. Do not append /chat/completions to the Base URL, and do not enter the site homepage, the billing page, or a model page URL. The existing https://bazaarlink.ai/api/v1 still works as a compatible endpoint, so you do not need to rebuild your key just to change the URL.
Step 3: Choose a Model and Send a Short Test Message
If a model list is available, choose from it; if fetching the list fails, you can type the full model ID manually. Confirm your account's free eligibility and remaining quota for the current period, then test the connection with auto:free. To specify a free DeepSeek variant from the current catalog, the ID is:
deepseek/deepseek-v4-flash-0731free:free
Clicking Test Message actually sends a request, which may count toward your free quota or paid usage. Do not click it ten times in a row while troubleshooting. Bypass API status check only suits the case where "the test message succeeded but the frontend status check still warns"; it cannot fix an invalid key, a permissions issue, or a quota problem.
For these connection options and limits, see the SillyTavern official custom endpoint documentation. OpenAI compatibility does not mean every tool, sampler, or extension is compatible.
Check the Free Quota First: "Free" Does Not Mean No Charges
Based on this check of the free models page and the public model catalog, the free items include Qwen3.7 Flash and the DeepSeek variant above. auto:free chooses among the free candidates available at the time and does not guarantee it will always return DeepSeek.
The public baseline is 10 requests per minute and 50 requests per day, and account tiers may have different multipliers. Your actual tier and whether you can use the free tier depend on your current balance, subscription, and service settings, not simply on whether you have topped up in the past. Account-specific restrictions may also make the free tier unavailable.
Successful calls within the free quota that go through the free route are not charged model fees. Beyond the quota, accounts that meet the paid fallback conditions may be moved to paid routing; otherwise the request returns 429. If you only want to test with free usage, check the errors and usage records first, and do not interpret auto:free as guaranteed zero billing. A model whose name only contains the word "free" but has no actual zero-price eligibility should not be assumed to be free either.
Before Choosing a Roleplay Model, Rule Out Settings-Caused Forgetfulness
When a character suddenly forgets a name, switches to English, or ignores the worldbuilding, the cause is not necessarily that "the model is not strong enough." Before switching to a paid model, check these four things:
- Whether the character card, system prompt, or preset contains conflicting instructions.
- Whether recent messages and key settings were removed from this request because of the context limit.
- Whether the relevant World Info entries were activated and whether they exceed their own token budget.
- Whether the output limit is too low, cutting off an otherwise complete reply midway.
World Info is not sent in full on every turn without condition. Official mechanics insert entries based on how they are triggered and their budget. Adding a setting does not guarantee the model will follow it; see the World Info official documentation.
After ruling out settings problems, fix the same character card, preset, context budget, output limit, and opening and compare a few candidates. Recording "which setting was missed, how many retries it took, and the total cost" will tell you more about whether a model suits you than a writing-quality ranking built without the same conditions.
The current candidates in this site's catalog include anthropic/claude-sonnet-5, google/gemini-3.1-pro-preview, and deepseek/deepseek-v4-pro. These are model names you can look up; they are not a tested winners list for roleplay in this article. For current unit prices and context sizes, see the model pages.
Estimating Costs: Roleplay Is Not the Same as Ordinary Chat
The input for one turn is more than the sentence you just typed. It may include the character settings, the preset, active World Info entries, the chat history sent, and content added by extensions. What to count is the input actually sent this time and the output the model generates, not how many sentences appear on screen.
Cost US$ = input tokens ÷ 1,000,000 × input price
+ output tokens ÷ 1,000,000 × output price
Worked Example for 100 Conversations
Using the Claude Sonnet 5 customer prices from this site's public catalog, with input at US$2 per million and output at US$10 per million, assume 100 requests where each request averages 10,000 input tokens and 500 output tokens, with no additional fees or cache discounts:
| Item | Calculation | Cost |
|---|---|---|
| Input | 100 × 10,000 = 1 million tokens | US$2.00 |
| Output | 100 × 500 = 50,000 tokens | US$0.50 |
| Total | Input + output | US$2.50 |
If you use US$1 = NT$31.5 for illustration, that is NT$78.75. This is a recalculable example, not "a typical player's average for one evening" or this site's checkout exchange rate. For the same 100 requests, if output averages 1,500 tokens instead, the total becomes US$3.50. Counting only input at US$2 would miss the output cost.
The prices come from the customer-price snapshot in the public model API on this site, not from official costs substituted for this site's billing. Your actual total is still determined by your own usage records. If reasoning or other billable items apply, check those too.
Control What Is Sent Before Considering Cache
Context Size can limit how much is retained, but a lower number is not always better: what gets cut may be exactly the basis for the character remembering something. When configuring, base the setting on the context capacity of the model you chose, reserve space for the generated output, and then look at how much the character, World Info, and history each take up.
Cache discounts do not apply automatically just because you paste in a character card. Check whether the actual model and route support them, whether the prefix is consistent, and whether cached usage actually appears on your bill. Record several real requests before judging how much you saved, rather than summarizing all settings with a single "switching models is five to ten times cheaper."
The Key Is Pasted but Still No Response: Follow the Error, Don't Switch Models First
| What you see | What to check first |
|---|---|
| URL or connection error | Base URL, whether a generation path was appended twice, and whether the frontend host can reach the outside network |
| 401 / 403 | The key and account permissions; a 403 can also come from model or content restrictions, so read the reason |
| 429 | Per-minute / per-day quota or other limits; pressing Bypass will not make it go away |
| model not found | Copy the full ID from the live catalog; do not type only the product display name |
| Replies are truncated or ignore settings | Output limit, context, preset, and World Info budget |
When asking for help, include your frontend version, model ID, error code, and a de-identified screenshot of your settings. Do not include your full key or private chat content. This site and the model channels may have usage policies or content restrictions; this article does not provide methods for getting around them.
Phones, Cloud Taverns, and Data Safety
A phone can connect to SillyTavern running on your own computer or a cloud host, but the host must keep running, be reachable over the network, and have appropriate login and connection protection. Do not expose an unauthenticated service to the internet just so your phone can connect. A cloud setup also adds hosting costs, backups, updates, and responsibility for storing your key. It is not just a question of "can I still play when my computer is off."
If you only want a phone frontend and are not committed to the original interface, you can compare options such as the RisuAI setup guide. Before switching frontends, confirm which APIs it supports, its character card formats, and how it stores data and keys. Do not assume that "it accepts a URL" means every extension is compatible.
Taiwan Platform Trade-offs: Choose Based on the Service You Actually Need
This site offers NT-dollar payments, API keys, and a multi-model catalog, and transactions on this site can be handled through the checkout process for invoices and Tax IDs. This helps people who need local payment or Chinese-language support. It does not mean overseas payments always fail, or that overseas receipts can never be reimbursed. Companies and schools should still have their accountant confirm the vouchers and budget rules; see the Taiwan Expense Reporting Guide.
Start with a short test using the free quota you are eligible for, then confirm usage and budget. Being able to switch models, understanding the cost, and keeping your character settings from being trimmed is what this tutorial is meant to help you accomplish.
FAQ
Which Base URL should I enter to connect SillyTavern to BazaarLink?
Set the API type to Chat Completion and the source to Custom (OpenAI-compatible), then enter https://api.bazaarlink.ai/v1 as the Base URL. Do not append /chat/completions. The existing https://bazaarlink.ai/api/v1 is still a compatible endpoint. Get your key and model IDs from your account on this site and the live catalog.
Is auto:free always completely free of charge?
Successful calls within the free quota that go through the free route are not charged model fees, but zero billing is not unconditionally guaranteed. The public baseline is 10 RPM / 50 RPD, and actual eligibility and multipliers depend on your current balance, subscription, and settings. Beyond the quota, accounts that meet the paid fallback conditions may be moved to paid routing; otherwise the request returns 429. Check your quota, errors, and usage records.
If the character forgets the worldbuilding, do I have to switch to a paid model?
Not necessarily. First check whether the character card and preset contradict each other, whether the context limit is cutting off important history, whether the relevant World Info entries are active or exceed their budget, and whether the output limit is truncating replies. After ruling out settings problems, compare candidate models using the same character card, preset, context budget, output limit, and opening, so that differences in settings are not mistaken for differences in model ability.
How is the API cost of one tavern chat calculated?
Multiply the actual input and output tokens by their respective prices per million tokens, then add the two. The character settings, active World Info entries, and the chat history sent may all count toward input. This article gives an explicit example: Sonnet 5 customer price is US$2 / US$10, with 10,000 input and 500 output tokens per call, for a total of US$2.50 over 100 calls. This is not a player average or a checkout commitment.
Can Bypass API status check fix a failed connection?
It only suits cases where the endpoint can already produce a reply but the frontend status check still warns. It does not fix an invalid key, permissions, a wrong model ID, or quota limits. Read the original error first, then decide whether to enable it, so you do not send repeated test requests that generate more usage.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API