BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-09-15 · · Author:BazaarLink · Prompt Caching · Prompt cache · AI API discounts · OpenAI · Claude · Gemini · DeepSeek · AI API cost

Prompt Caching: Savings, Rules and BazaarLink Cache Prices

Prompt caching needs no model switch or batch wait, but OpenAI, Claude and Gemini differ in minimum length, lifetime and write fees, with BazaarLink prices.

Query date: 2026-09-15. The rules for the three providers are taken from each provider's official documentation as of that day; cache prices come from the BazaarLink Models catalog. Cache rules change often, so verify again before you act.

Prompt caching is currently the only API discount you can get without switching models, waiting 24 hours, or negotiating. The principle in one sentence: the opening section that repeats in every request — the system prompt, tool definitions, conversation history, the long document you paste in — is not paid at full price from the second time onward.

But it is not automatically on for everything. The minimum length, retention time, write fees and whether you need to add parameters all differ across the three providers, and getting any one of them wrong means zero discount with no error message. This article puts the three providers' rules in one table, gives you the actual cache prices in our catalog, and finally calculates how much you actually save.

The rules of the three providers in one table

OpenAI (GPT-5.6+)Anthropic (Claude)Google (Gemini)
How to enableOn by default, no parameter neededAdd cache_control: add it once at the top level = automatically placed on the last cacheable block; or place it manually block by blockModels 2.5 and above have implicit caching by default; there are also explicit cache objects
Minimum length1,024 tokensOpus 4.8 / Sonnet 5 1,024; Haiku 4.5 4,096; Fable 5.1 5123.x Flash and 3.1 Pro 4,096; 2.5 series 2,048
Hit conditionEntire prefix must be identical (including tool definitions and history)Content before the breakpoint must be identicalPrefix must be identical
Retention time30 minutes from the most recent read or writeDefault 5 minutes; optional 1 hourImplicit: automatic; explicit: custom TTL
Read price0.1× the input price0.1× the input price10% of the input price (3.8 Flash $0.075)
Write price1.25× the input price5-minute TTL 1.25×; 1-hour TTL 2×Implicit: not charged separately; explicit: storage billed hourly (3.8 Flash $0.50 per million tokens per hour, 3.1 Pro $4.50)
Breakpoint limit—Up to 4—
How to check whether a hit occurred in the responseusage.input_tokens_details.cached_tokensusage.cache_read_input_tokens, cache_creation_input_tokensCache fields in the response usage

Three of the most common pitfalls:

  1. Requests shorter than the threshold are not cached, and no error is raised. Claude's documentation states clearly that below the minimum length, even if you add cache_control, nothing is cached and no error is returned; you have to check the usage fields yourself. Haiku 4.5's threshold is 4,096, and many customer service bots' system prompts never reach that length.
  2. Changing one character in the middle of the prefix invalidates everything after it. A cache covers "identical from the beginning up to a certain point." Putting things that change (dates, user names) at the start of the system prompt means a new cache is written every time — and writes cost 1.25×, which is more expensive than not caching at all.
  3. Writes are not free. If a prefix appears only once, you pay 1.25×; it only starts to pay off when it appears a second time. Claude's 1-hour TTL write costs 2×, so it needs to repeat more times to be worthwhile.

When you call through BazaarLink, cache reads are billed according to the table below, and the matching rules are the same as direct purchase (same identical prefix, same minimum length).

ModelInputCache readRead = % of inputCache write
GPT-5.6 Sol$2.00$0.2010%$2.50
GPT-5.6 Terra$2.00$0.2010%$2.50
GPT-5.6 Luna$0.20$0.02010%$0.25
Claude Opus 4.8$5.00$0.5010%$6.25
Claude Sonnet 5$2.00$0.2010%$2.50
Claude Haiku 4.5$1.00$0.1010%$1.25
Gemini 3.8 Flash$0.75$0.07510%Not listed in catalog
DeepSeek V4 Flash$0.20$0.03015%Not listed in catalog
DeepSeek V4 Pro$2.40$0.146%Not listed in catalog

Does the platform actually apply the discount? We reviewed the usage records billed by the platform over the past 7 days: 246 OpenAI-series requests, 2,600 DeepSeek requests and 74 Gemini requests hit the cache, and the cache discounts were actually posted. All cache-hit requests in the Claude series during these 7 days were bring-your-own-key (BYOK) usage — that money is paid by the customer to Anthropic and does not pass through our billing, so the platform currently has no actual Claude discount sample; the Claude cache prices in the table above are catalog prices, and the billing mechanism is the same as the other three. We state this as it is and do not work around it.

How much you actually save: three scenarios calculated

Scenario 1: Customer service bot, 60% of input hits the cache (Claude Sonnet 5)

Monthly input 13.5M tokens, output 6.75M tokens. The system prompt plus the first few rounds of history make up about 60% of input.

  • No caching: 13.5 × $2.00 + 6.75 × $10.00 = $94.5
  • With caching: 13.5 × (40% × $2.00 + 60% × $0.20) + 6.75 × $10.00 = $79.9
  • Savings 15%

Scenario 2: Document batch, long document asked many questions, 90% of input hits the cache (GPT-5.6 Luna)

The same 6,500-token document is asked 10 questions; the document itself hits the cache from the second question onward. Monthly input 19.5M, output 2.4M.

  • No caching: $6.78; with caching: $3.62; savings 47%

Scenario 3: The same document batch, switching to DeepSeek V4 Flash

  • No caching: $4.86; with caching: $1.88; savings 61%

All three scenarios show the same thing: how much you save depends on "how much of the bill is input" × "how much of the input is repeated." Output-dominated workloads (generating long text, reasoning models) get little help from caching; input-dominated and highly repetitive workloads (RAG, many questions on long documents, customer service with a fixed system prompt) are where it works best.

How to order the prefix

The order is always: put what changes least at the front.

  1. Tool definitions (tool schema)
  2. The fixed part of the system prompt
  3. Long documents / knowledge base content
  4. Conversation history (older turns first)
  5. Put things that change at the end: the current date, the user's name, the question in this turn

Claude allows up to 4 breakpoints among items 1 through 4, so sections that change at different frequencies can each be cached; OpenAI and Gemini have no breakpoints and rely on the entire prefix aligning naturally.

Nothing. Send requests in the OpenAI-compatible format as usual; cache hit detection and billing are handled upstream and in our billing layer. The usage field in the response returns the cached token count, so you can see how many tokens each request hit in your usage records. The full price table is in the cross-provider price comparison, which also covers other discount mechanisms (batch, off-peak, usage rebates) — we do not offer a Batch API; if you need one, you would need to purchase it directly from the provider.

FAQ

What is prompt caching?

Cache the repeated opening section that appears in each request (system prompt, tool definitions, conversation history, long documents), so from the second time onward it is billed at the cache price instead of the full price. OpenAI and Gemini have it on by default, and Claude requires cache_control. The read price is usually 10% of the input price; OpenAI and Claude charge 1.25× to write, so the prefix must repeat at least twice to pay off.

How much can prompt caching save?

It depends on how much of the bill is input and how much of the input is repeated. A customer service bot using Claude Sonnet 5 with a 60% input hit rate saves about 15%; long-document multi-question answering using GPT-5.6 Luna with a 90% hit rate saves over 60% of the input fees. Output is not affected by caching, so output-dominated workloads save almost nothing.

Why do I get no discount even after adding cache_control?

The most common cause is a prefix shorter than the minimum length: GPT-5.6 and Claude Sonnet 5 / Opus 4.8 require 1,024 tokens, while Claude Haiku 4.5 and Gemini 3.x require 4,096. Below the threshold, nothing is cached and no error is returned; check the cache fields in the response usage. The second most common cause is content in the middle of the prefix that changes (dates, user names), which creates a new cache every time.

How long is the cache kept?

OpenAI GPT-5.6 and later: 30 minutes from the most recent read or write. Claude: 5 minutes by default, optionally 1 hour (writes become 2×). Gemini: implicit caching is managed automatically, and explicit caching has a custom TTL but storage is billed hourly.

Does the cache still work when called through BazaarLink?

Yes, and no code changes are needed. Cache reads are billed at the cache price in the catalog (for most GPT, Claude and Gemini models, 10% of the input price; DeepSeek V4 Flash 15%). In the past 7 days, OpenAI, DeepSeek and Gemini requests billed by the platform all posted actual cache discounts. For Claude, all cache-hit requests during this period were BYOK usage, so the platform has no actual sample yet, but the catalog price and billing mechanism are the same as the other three.

What is the difference between Gemini explicit and implicit caching?

Implicit caching is enabled by default for models 2.5 and above; when the prefix reaches the minimum length (4,096 tokens for 3.x Flash and 3.1 Pro), it hits automatically with no extra charge. Explicit caching means you create a cache object manually and specify a TTL; cached input is also 10%, but storage is billed hourly (3.8 Flash $0.50 per million tokens per hour, 3.1 Pro $4.50).

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
Claude Haiku 5.5 · DeepSeek V4.1 Flash · GLM-5.3-Flash · MiMo-V2.6-Flash · Model Comparison · Pricing
Haiku 5.5 vs DeepSeek, GLM, MiMo Flash: Price and AA Index
Claude · Haiku 5.5 · Haiku 4.5 · Model Comparison · Pricing
Haiku 4.5 or 5.5? $0.10/$0.50 Specs and NT$ Cost Estimates
Gemini 4 Argon · Vending-Bench 2 · AI agent · Andon Labs · agent safety
Gemini 4 Argon: #3 on Vending-Bench 2, accused of lying
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.