BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-09-05 · · Author:BazaarLink · Tavern AI · Tavern · SillyTavern · AI roleplay · beginner guide · free model · Taiwan

What Is Tavern AI? Free Trial, Mobile Play, API Costs

SillyTavern is a front end and needs its own model. Get one reply first, then add a character card. Covers free tiers, limits, paid fallback and costs.

Verification date: 2026-09-17 (Taipei time). This article updates the official setup docs, free models and quota notes, but no fresh installation or paid conversation test was run. Get a short message to reply successfully first, then add character cards and world info. Installation, network and account eligibility differ, so we cannot guarantee that everyone finishes in ten minutes.

"Tavern AI" usually refers to SillyTavern: a front end that lets you arrange characters, prompts and chats yourself. If it is installed but not talking, you most likely have not connected a model service that generates replies, rather than missing a character pack.

You need two things: a usable chat front end, and a model service it can connect to. An API key is the access credential you get from that service, not a universal free pass for any app.

What is Tavern AI? How does it differ from ChatGPT and Character.AI?

What to compareIntegrated platforms such as ChatGPT / Character.AISillyTavern
Model connectionThe platform arranges it under its own plansYou choose a compatible API, local model or other source
Character and prompt settingsDepends on platform features and permissionsYou can manage character cards, presets, world info and other settings
Chat history and backgroundDepends on the platform's actual mechanismsYou can adjust what gets sent, but it is still bound by model and context limits
CostFree or paid plans, per each platform's rulesFront end software, model and hosting costs should be viewed separately

SillyTavern has many adjustable options, which also means you are responsible for settings, updates and keeping your keys. More features do not make it better for beginners: if you only want to chat on a website and do not want to manage character settings or APIs, an integrated platform may be easier.

Where do the models come from? Trade-offs of three routes

Apply for an API directly from a model service

Handle it according to each service's region, account verification, free credits and payment terms. Do not treat "Taiwan credit cards will always be rejected" as a rule that holds for every provider. An API account may not share quota with a chat website subscription, so clarify this before applying.

Community compute: AI Horde

SillyTavern has an AI Horde connection option. Guest priority is low by default, so waits may be long. It is powered by volunteer compute. Do not send real names, emails, private conversations or other personal data. The official docs also warn that not every worker can be guaranteed not to log content; see AI Horde official settings and risk notes. Being free and being suitable for sensitive data are two different questions.

Compatible APIs on Taiwanese platforms

For example, this site offers API keys, payment in New Taiwan dollars and a free model quota. It may suit you to first see whether you like this way of playing, but check account eligibility and limits first. The free quota is not unlimited, and not every request is necessarily zero-cost.

Desktop: get the first reply before importing a full character card

If you have not installed it yet, first follow the SillyTavern official installation page for your system. The steps below are the connection setup once SillyTavern already opens; they are not a full installation process.

  1. Create an API key on this site's keys page and confirm your current account eligibility and quota. Store the key in a safe place; do not put it in a public character card.
  2. In API Connections, choose Chat Completion, and set the source to Custom (OpenAI-compatible).
  3. Fill in the Base URL as https://api.bazaarlink.ai/v1. Do not add /chat/completions. Enter this site's key.
  4. Choose a model from the live list, or type the full ID manually. If you currently qualify for free use, you can test with auto:free.
  5. Send one short Test Message. After it succeeds, import one character card to test the tone. Do not start by loading many extensions, long world info and complex presets.

This connection method follows the SillyTavern official custom endpoint guide. Existing settings that use https://bazaarlink.ai/api/v1 still work. For field details and troubleshooting, see SillyTavern Taiwan API setup.

Do not treat "the API gave its first reply" as full compatibility. Character cards, tools, extensions and prompt formats still need to be checked one by one. That is the reason to start with short messages.

Free models and limits: when might it switch to paid?

Checking the public model catalog and the free page this time, the free items include Qwen3.7 Flash and the DeepSeek variant with ID deepseek/deepseek-v4-flash-0731free:free. auto:free picks a currently available free candidate; it does not guarantee the same model every time.

The public baseline is 10 requests per minute and 50 requests per day, and tiers may carry multipliers. Actual eligibility depends on the current balance, subscription and service settings. It is not "you topped up once and will stay at the same tier forever."

Successful requests that route to free models within quota are not charged model fees. Beyond quota, accounts that meet the conditions for paid fallback may be switched to paid; accounts that do not will receive 429. Check your own limits and usage. Tests, regenerations and extensions can each send API requests that count toward the limit, so do not treat one "turn" as a single API call.

Character forgets things? Check settings before buying a pricier model

First confirm whether important settings are still in the context sent this time, whether the character card and preset conflict, whether world info is activated, and whether the output is cut off by the limit. Seeing a piece of history in the chat window does not mean this request actually sent it to the model.

After ruling these out, fix the same card, preset, context and output budget and compare candidates. Evaluate what the character remembers, whether it keeps its tone, how many retries are needed and the total cost, not just the "most expensive model."

World info is a tool for dynamically adding background, controlled by activation methods and token budgets. It is not unlimited memory. For the mechanism, see the official World Info docs.

No computer? First distinguish original Tavern from mobile front ends

You can use a phone, but there is more than one route:

Way to playWhat must keep runningWhat to confirm first
Original on a computer or host, phone browser connects inHost kept running with a reachable networkLogin protection, connection security, backups; do not expose an unverified service publicly
Self-host the original on AndroidThe phone's runtime environment and background servicesFollow the official Android installation docs; evaluate update and background-stop issues
Use an independent mobile front endThat app's account and settingsCustom API, character card format, how keys and chats are stored, and its own costs

For example, MiniTavern's official Android route guide describes connecting a custom OpenAI-compatible endpoint, and it also states clearly that it is not SillyTavern's official mobile version. Supporting a given API does not mean supporting all of the original's character cards, plugins and sampling settings. Confirm downloads and versions from that product's current official entry point.

If you do not want to install or maintain a host, you can also compare the RisuAI setup guide. Whether this site's key can be used in a given front end depends on its API compatibility and service policy. Do not scatter the same key across unfamiliar apps. When needed, create a separate key for each front end, for easier revocation and usage tracking.

Got a key but nothing happens? Read the error instead of pressing Connect repeatedly

First check the Base URL, the full model ID and the key source. For 401 or 403, read the reason given for verification, permissions or restrictions. For 429, check quota and limits first. If a model does not appear in the list, check the live catalog and enter it manually rather than typing only the product's Chinese name.

Bypass API status check is only for when the endpoint already replies successfully but the front end's status check still warns. It is not a button that makes errors, blocks or quota disappear. Before sharing a screenshot for help, hide your key, email and private chats. Content and usage policies still apply, and switching front ends does not remove restrictions.

After switching to a paid model, how do you estimate costs?

Input includes the settings and history actually sent, and output is what the model generates this time. Both affect the cost. Do not count only the short lines you typed yourself.

Model cost per request = actual input tokens × input unit price
                       + actual output tokens × output unit price

For example, the full setup article uses this site's Sonnet 5 customer price for an explicit estimate: 100 requests, averaging 10,000 actual input and 500 output tokens each, total US$2.50. This is not the average player's usage. Longer output, more history, background requests or other billing items will change the total.

When first switching to a paid model, cap the output and set an acceptable budget. Check usage after chatting a few times, then decide whether to lengthen the context. A New Taiwan dollar conversion example cannot replace the actual checkout exchange rate.

Should you use "cloud Tavern"?

Cloud Tavern usually means hosting the front end on a cloud server. The advantage is not depending on your own computer being on. The costs, besides host fees, include updates, access protection, key storage, chat backups and failure handling. It is not just one extra rental payment with nothing else to manage.

If you are just starting, first choose a front end where you understand how data and costs flow. Do not rush into deployment or the most expensive model. First confirm that you like this style of chat, then invest maintenance time and quota.

FAQ

Is Tavern AI free?

The front end software, the model service and hosting costs should be considered separately. Requests that succeed within this site's quota and go through the free model route are not charged model fees, but they are subject to account eligibility and request limits. Beyond the quota, eligible accounts may be moved to paid; otherwise the request returns 429. Do not assume that a free front end or a model name containing free means everything costs nothing.

Why does SillyTavern not reply after installation?

It is the front end that arranges chats, characters and prompts, and it needs a model service it can connect to. Check the API type, source, Base URL, key and full model ID, then send one short Test Message. After that succeeds, add character cards, world info and presets; do not mix in many variables at the start.

Can beginners definitely start in ten minutes?

No guarantee. Installation, network, version, account eligibility and API limits all affect how long it takes. The connection steps here assume SillyTavern already opens; this is not a full installation guide. Get a short message reply first, then test one character card.

Can I play the original Tavern on a phone?

You can connect to the original on your own host, or self-host it following the official Android documentation; the host and network must be usable and access-protected. An independent mobile front end is a different product, such as MiniTavern, which is not SillyTavern's official mobile version. API, character cards, plugins and data storage each need to be confirmed; full compatibility is not guaranteed.

If a character forgets names, does that mean the free model is not good enough?

Not necessarily. First check whether important history was sent, whether the context was trimmed, whether the preset conflicts, whether world info is activated, and whether the output was truncated. After ruling out settings problems, fix the same card and budget and compare models. A chat window still showing history does not mean this request included that history.

How is the cost calculated after switching to a paid model?

Input includes the character settings, world info and history sent in the request, and output is charged at its unit price too. The full setup article gives a worked example: for Sonnet 5 at the customer price of US$2 / US$10 per million, 100 requests averaging 10,000 input and 500 output tokens total US$2.50. This is not the average or a guaranteed bill for every player.

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
Claude Code · Claude Code update · changelog · mods · permission rules
Claude Code 2.1.289: 27 Changes, Permission-Rule Fixes and Mods Stability
Claude Code · Claude Code mods · plugins · TypeScript · AI coding tools
Claude Code Mods: Customize Behavior and UI with TypeScript
web search API · AI agent · LLM · Hermes Agent
Hermes Agent Web Search: Use a BazaarLink Model with :online
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.