2026 AI API Pricing in TWD: Claude, GPT-6, Gemini, DeepSeek
2026 AI API price comparison: Claude, GPT-6, Gemini, DeepSeek per-million-token USD/TWD, with caching, batch, usage rebates and a cost estimate. Prices dated.
This table lists the multi-provider API unit prices (USD per million tokens) from BazaarLink's public model catalog. Each model's direct official price and its batch, promotion, and peak/off-peak conditions may differ. The TWD figures are budget estimates, using the Bank of Taiwan USD spot selling rate of NT$31.775 as of 2026-09-23. Catalog data and official price verification date: 2026-09-24. You can verify against the public model catalog and each provider's official price list.
Cross-provider price comparison (USD / NT$, per million tokens)
| Model | Provider | Positioning | Input | Output | NT$ Input | NT$ Output |
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Top tier | $10.00 | $50.00 | NT$ 10,009 | NT$ 50,046 |
| GPT-5.6 Sol | OpenAI | Flagship workhorse | $2.00 | $10.00 | NT$ 2,002 | NT$ 10,009 |
| GPT-5.6 Terra | OpenAI | Workhorse | $2.00 | $12.00 | NT$ 2,002 | NT$ 12,011 |
| GPT-5.6 Luna | OpenAI | Lightweight | $0.20 | $1.20 | NT$ 200 | NT$ 1,207 |
| Claude Opus 4.8 | Anthropic | Top tier | $5.00 | $25.00 | NT$ 5,020 | NT$ 25,039 |
| Claude Sonnet 5 | Anthropic | Workhorse | $2.00 | $10.00 | NT$ 2,002 | NT$ 10,009 |
| Claude Haiku 4.5 | Anthropic | Lightweight | $1.00 | $5.00 | NT$ 1,017 | NT$ 5,020 |
| Gemini 3.1 Pro Preview | Long-context workhorse | $2.00 | $12.00 | NT$ 2,002 | NT$ 12,011 | |
| Gemini 3.8 Flash | Workhorse (promotional price through 2026-12-31) | $0.75 | $3.75 | NT$ 763 | NT$ 3,749 | |
| Gemini 3.5 Flash-Lite | Lightweight | $0.30 | $2.50 | NT$ 299 | NT$ 2,510 | |
| Gemini 2.5 Flash-Lite | Ultra-lightweight | $0.10 | $0.40 | NT$ 102 | NT$ 413 | |
| DeepSeek V4 Pro | DeepSeek | Workhorse | $2.40 | $4.80 | NT$ 2,415 | NT$ 4,798 |
| DeepSeek V4 Flash | DeepSeek | Lightweight | $0.20 | $0.40 | NT$ 200 | NT$ 413 |
| Qwen3.7 Flash | Qwen | Free-tier model | $0.030 | $0.13 | NT$ 29 | NT$ 130 |
Three things are visible directly from the table: the top tier (GPT-6 Astra at $10 / $50) and the cheapest paid model (DeepSeek V4 Flash at $0.20 / $0.40) differ by 50x and 125x; DeepSeek V4 Pro's $2.40 input is workhorse-level, but its $4.80 output is less than half of Sonnet's, so for generation-heavy uses the results differ greatly; Gemini 3.8 Flash's $0.75 / $3.75 is the price officially marked as valid through 2026-12-31, so keep that in mind when budgeting.
Addendum (2026-10-08): The "cheapest paid model" above is a comparison based on the output unit price of DeepSeek V4 Flash in the table ($0.40). Models with lower input prices are Gemini 2.5 Flash-Lite in the table ($0.10) and Claude Haiku 5.5, listed on 2026-10-07 and not yet in the table (input $0.10 / output $0.50; its output price is still higher than DeepSeek V4 Flash). Neither changes the 50x and 125x figures, which use DeepSeek V4 Flash as the denominator.
OpenAI API Price: GPT-6 pricing, API keys and first call
OpenAI API pricing is calculated by model and token usage. The official price list as of 2026-09-24 shows GPT-6 Sol standard input/output at US$2 / US$10, cache reads at US$0.20, and cache writes at US$2.50 per million tokens; long-context requests have separate price conditions. OpenAI Platform API keys are created from a project and billed separately from ChatGPT subscriptions.
Applying for an OpenAI API key
- Log in to OpenAI Platform and select your organization and project.
- Open Settings → Project → API Keys, click Create new secret key, name it, and set permissions according to its purpose.
- Put it in the OPENAI_API_KEY environment variable. Do not write it into web frontend code or the repository.
Official project key management and API quickstart documentation, verified 2026-09-24: Projects and API keys, API quickstart.
First call to the OpenAI official and compatible APIs
Official Responses API:
export OPENAI_API_KEY="your OpenAI project key"
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-6-sol","input":"Explain API token pricing in one sentence in Traditional Chinese."}'
Using the OpenAI-compatible Chat Completions format:
export BAZAARLINK_API_KEY="your BazaarLink key"
curl https://api.bazaarlink.ai/v1/chat/completions \
-H "Authorization: Bearer $BAZAARLINK_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Explain API token pricing in one sentence in Traditional Chinese."}]}'
OpenAI model price table (2026-09-24)
The following are standard prices per million text tokens. The NT$ columns are converted using the Bank of Taiwan USD spot selling rate of 31.775 as of 2026-09-23, for budget estimates only.
| OpenAI model | Input USD | Cache read USD | Cache write USD | Output USD | Input est. NT$ | Output est. NT$ |
|---|---|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 | approx. 317.75 | approx. 1,588.75 |
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 | approx. 63.55 | approx. 317.75 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 | approx. 3.18 | approx. 15.89 |
This is the Standard price for short and typical context lengths. Long-context requests for models such as GPT-6 Sol, other service tiers, and tool calls are billed separately under official pricing. For BazaarLink's current prices, you can also check the OpenAI GPT-6 Sol model page, public model IDs and unit prices, and cache and discount fields. Sources: OpenAI official API pricing, GPT-6 Sol detailed pricing, Bank of Taiwan exchange rate, all verified on 2026-09-24.
The cross-provider comparison table that follows retains the original 2026-09-15 snapshot date. The OpenAI section above has been separately updated to 2026-09-24. For other providers' prices, check the date marked on the original table and each provider's official source. Further reading: OpenAI API Taiwan purchasing and billing, Claude API, Gemini API fees and TWD estimate, Gemini API sign-up.
Query date: 2026-09-15. Unit prices in the table are the customer unit prices publicly listed on the BazaarLink Models page, which match each provider's official list prices. Cache prices are taken from the same catalog. The exchange rate is 1 USD = NT$ 31.5, consistent with other price articles on this site. Providers update their official pages often, so check again before acting.
Searches for "AI API price", "cost", "discounts" and "promotions" turn up the same set of articles, because in Taiwan these four things are one question: how much each million tokens actually costs, how much my use case will cost per month, and what methods can make it cheaper. This article puts four mainstream models in one table, runs estimates on the same set of scenarios, and then splits "discounts" into four mechanisms that really exist, including the two that BazaarLink does not offer.
Summary first:
- Flagship models are priced the same: GPT-5.6 Sol and Claude Sonnet 5 are both $2 / $10. Price alone cannot decide; choose by quality and latency.
- The output unit price drives the bill: most models' output costs 5x their input, so for generation-heavy workloads look at the output columns on the right first.
- There are four kinds of "discount", and two are only available when buying direct: cache reads (BazaarLink has them; for most models they are 10% of the input price), promotional prices (yes, mirroring official), Batch API 50% off (we do not have it), and DeepSeek off-peak half price (only available when buying direct). One more exists only with us: usage rebates.
- The cheapest and the most expensive differ by 50–125x: model choice affects cost more than any discount.
Why the bill is higher than you calculated: look at output first, then caching
| Model | Output / input multiple |
|---|---|
| GPT-6 Astra | 5.0x |
| GPT-5.6 Sol | 5.0x |
| GPT-5.6 Terra | 6.0x |
| GPT-5.6 Luna | 6.0x |
| Claude Opus 4.8 | 5.0x |
| Claude Sonnet 5 | 5.0x |
| Claude Haiku 4.5 | 5.0x |
| Gemini 3.1 Pro Preview | 6.0x |
| Gemini 3.8 Flash | 5.0x |
| Gemini 3.5 Flash-Lite | 8.3x |
| Gemini 2.5 Flash-Lite | 4.0x |
| DeepSeek V4 Pro | 2.0x |
| DeepSeek V4 Flash | 2.0x |
Long document, short answer (summarization, classification, extraction) is input-dominated. Short question, long answer (article generation, code generation, reasoning models that output their thinking process) is output-dominated. First determine which type your workload is, then check the table. Otherwise you may choose a model that is "cheap on input but expensive on output".
Four discount mechanisms: which are only available when buying direct, which we have, and which only we have
| Mechanism | How it works | Direct from official | BazaarLink |
|---|---|---|---|
| Cache read price (prompt caching) | Repeated prefixes (system prompt, conversation history, long documents) are billed at the cache price from the second occurrence onward | Each provider offers it | Yes. For most models, cache read = 10% of the input price (see table below) |
| Promotional price | Official time-limited price | Yes | Yes, mirroring official (e.g., Gemini 3.8 Flash through 2026-12-31) |
| Batch API 50% off | Not real-time, returned within 24 hours, input and output halved | OpenAI, Anthropic and Google offer it | No |
| Off-peak half price | DeepSeek direct: all models at half price outside peak hours (peak = UTC Mon–Fri 01–04 and 06–10, i.e., Taipei 09–12 and 14–18) | DeepSeek offers it | No, we use fixed prices |
| Usage rebate | When monthly spending crosses a threshold, a fixed amount is issued; reached thresholds accumulate, and 10% is returned above $500 (DeepSeek 20%) | No | Only we offer it, rules |
| Reseller enterprise discount | Negotiated discount rate | — | Enterprise plans have it; the rate is not public and is not stated in this article |
If your workload is "not urgent, high volume, and can wait 24 hours", the Batch API is the largest discount, and it requires buying direct. We are not skipping past this.
Our cache prices (as a fraction of the input price)
| Model | Input | Cache read | Read as % of input | Cache write |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | 10% | $12.50 |
| GPT-5.6 Sol | $2.00 | $0.20 | 10% | $2.50 |
| GPT-5.6 Terra | $2.00 | $0.20 | 10% | $2.50 |
| GPT-5.6 Luna | $0.20 | $0.020 | 10% | $0.25 |
| Claude Opus 4.8 | $5.00 | $0.50 | 10% | $6.25 |
| Claude Sonnet 5 | $2.00 | $0.20 | 10% | $2.50 |
| Claude Haiku 4.5 | $1.00 | $0.10 | 10% | $1.25 |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | 10% | $0.38 |
| Gemini 3.8 Flash | $0.75 | $0.075 | 10% | — |
| Gemini 2.5 Flash-Lite | $0.10 | $0.010 | 10% | $0.083 |
| DeepSeek V4 Pro | $2.40 | $0.14 | 6% | — |
| DeepSeek V4 Flash | $0.20 | $0.030 | 15% | — |
For the GPT and Claude series, cache writes cost 1.25x the input price. Caching only pays off when the prefix repeats more than once.
Three scenarios, same assumptions, four providers side by side (monthly)
Assumes all runs are 30 days and caching is not counted. The caching effect is demonstrated separately in the next section.
Customer service bot: 500 conversations/day, 3 turns per conversation, 300 in / 150 out per turn
Monthly tokens: input 13.5 M, output 6.75 M.
| Model | Monthly USD | Monthly NT$ |
|---|---|---|
| DeepSeek V4 Flash | $5.4 | NT$ 5,402 |
| GPT-5.6 Luna | $10.8 | NT$ 10,804 |
| Gemini 3.5 Flash-Lite | $20.9 | NT$ 20,940 |
| Gemini 3.8 Flash | $35.4 | NT$ 35,461 |
| Claude Haiku 4.5 | $47.2 | NT$ 47,281 |
| Claude Sonnet 5 | $94.5 | NT$ 94,594 |
| GPT-5.6 Terra | $108 | NT$ 108,099 |
For the same task, the cheapest and the most expensive differ by 20x.
Document batch: 100 documents/day, each 6,500 tokens read and 800 tokens summarized
Monthly tokens: input 19.5 M, output 2.4 M.
| Model | Monthly USD | Monthly NT$ |
|---|---|---|
| Gemini 2.5 Flash-Lite | $2.9 | NT$ 2,923 |
| DeepSeek V4 Flash | $4.9 | NT$ 4,862 |
| GPT-5.6 Luna | $6.8 | NT$ 6,800 |
| Gemini 3.5 Flash-Lite | $11.8 | NT$ 11,852 |
| Gemini 3.8 Flash | $23.6 | NT$ 23,641 |
| Claude Haiku 4.5 | $31.5 | NT$ 31,521 |
For the same task, the cheapest and the most expensive differ by 11x.
Long document analysis: 50 documents/day, 50,000 tokens each, 2,000 output
Monthly tokens: input 75 M, output 3 M.
| Model | Monthly USD | Monthly NT$ |
|---|---|---|
| DeepSeek V4 Flash | $16.2 | NT$ 16,205 |
| Gemini 3.8 Flash | $67.5 | NT$ 67,554 |
| Claude Sonnet 5 | $180 | NT$ 180,164 |
| GPT-5.6 Sol | $180 | NT$ 180,164 |
| Gemini 3.1 Pro Preview | $186 | NT$ 186,170 |
| DeepSeek V4 Pro | $194 | NT$ 194,590 |
| Claude Opus 4.8 | $450 | NT$ 450,411 |
For the same task, the cheapest and the most expensive differ by 28x.
How much caching saves: using the customer service bot as an example
In each turn of the customer service bot's input, the system prompt and the first few turns of history repeat. Assume 60% of input tokens hit the cache and use Claude Sonnet 5:
- No caching: 13.5 × $2.00 + 6.75 × $10.00 = $94.5/month
- With caching: 13.5 × (40% × $2.00 + 60% × $0.20) + 6.75 × $10.00 = $79.9/month
- Savings of 15%. The output half is unchanged, so the overall percentage saved is always less than the caching discount itself.
For caching details (which models support it, how prefixes are counted, minimum lengths), see a separate article: What is prompt caching and how much it saves.
How to choose: by use case, not by brand
| Use case | Look at first | Suggested starting point |
|---|---|---|
| Classification, extraction, summarization (input-dominated) | Input unit price | Gemini 3.5 Flash-Lite, DeepSeek V4 Flash, GPT-5.6 Luna |
| Customer service, translation, moderate complexity | Output unit price + caching | Gemini 3.8 Flash, Claude Haiku 4.5 |
| Long-form generation, code | Output unit price | DeepSeek V4 Pro (output $4.80), Sonnet 5 |
| Long documents, long context | Input unit price + context length | Gemini 3.1 Pro Preview (input $2, 1M context) |
| Highest quality, cost no object | — | GPT-6 Astra, Claude Opus 4.8 |
| Get connected and validate first | Free tier | Qwen3.7 Flash (free tier, 10 per minute, 50 per day) |
Taiwan dollars, invoices, rebates: three things beyond the bill
- TWD pricing and unified invoices: Buying direct gives you a USD overseas receipt. BazaarLink charges in TWD top-ups, and eligible payment channels can issue unified invoices (personal or business tax ID). For the expense reporting process, see this article.
- Usage rebate: when monthly spending on the same provider (OpenAI GPT / Gemini / DeepSeek) crosses $20 / $50 / $100 / $200 / $500, you receive $2 / $3 / $5 / $10 / $30 respectively (DeepSeek: $4 / $6 / $10 / $20 / $60). Above $500, 10% is returned again (DeepSeek 20%). Nothing below $20; it becomes stable only from $500 upward, full rules.
- Key spending cap: each API key can have a spending limit. Once reached, requests are blocked; it is not just an alert.
Standalone TWD pricing articles by provider
GPT-5 series, Claude, Gemini (including free credits), DeepSeek. For a summary of free credits, see this article.
Explore current model discounts and learn how usage rebates add credit to your balance.
View discounts and usage rebatesFAQ
Which AI API is the cheapest? Which matters more when comparing, input or output?
There is no single lowest-price model that suits every workload. Compare the input and output unit prices per million tokens separately, along with caching and service tier. For short questions with long answers, look at the output price; for long documents with heavy input, look at the input price. This table compares Claude, GPT, Gemini, DeepSeek and other providers' models in the same unit, and notes sources and dates. Price verification date: 2026-09-24.
What discounts, prompt caching or usage rebates are available for AI APIs?
Each program has different terms: some providers' Batch tiers have lower unit prices; cache reads are billed only on cache-hit tokens; DeepSeek's official direct prices are listed separately for peak and off-peak hours. The BazaarLink usage rebate credits balance on eligible paid usage and does not change the unit price of any single API call. Caching savings can be estimated as 'cache-hit tokens x (regular input price - cache read price)', then adding back cache write and storage costs. If the prefix rarely repeats, caching may not pay off. Check each condition against the official prices and public rules cited in this article. Date: 2026-09-24.
Are BazaarLink prices more expensive than buying direct?
The table shows BazaarLink public model catalog unit prices captured on 2026-09-24. Each provider's official direct prices vary by model, context length, Batch, promotions, and peak or off-peak hours. Before comparing, check the same model ID, service tier, input and output token categories, and that day's price list. Do not compare based on a single aggregate figure.
What is the GPT-6 Sol API price?
As of 2026-09-24, OpenAI's official standard price is US$2 per million input tokens and US$10 per million output tokens, with cache reads at US$0.20 and cache writes at US$2.50. Long context and other service tiers are billed separately.
What exchange rate is used for the TWD conversion in the OpenAI API Price table?
The TWD examples in this article's OpenAI price table use the Bank of Taiwan USD spot selling rate of NT$31.775 on 2026-09-23, for budget estimates only. Actual transaction rates may differ. Model price verification date: 2026-09-24.
Is an OpenAI API key billed the same as ChatGPT Plus?
No. OpenAI manages ChatGPT subscriptions and API Platform usage separately. API key usage fees are calculated by model, tokens and official pricing. Verification date: 2026-09-24.
Where can I see the OpenAI models BazaarLink can call?
Check the public model list for model IDs and prices. Cache and discount fields are listed separately in the public price catalog. Examples and links are in the OpenAI pricing section of this article.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API