BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-10-08 · · Author:BazaarLink · Claude Haiku 5.5 · DeepSeek V4.1 Flash · GLM-5.3-Flash · MiMo-V2.6-Flash · Model Comparison · Pricing

Haiku 5.5 vs DeepSeek, GLM, MiMo Flash: Price and AA Index

Haiku 5.5 costs $0.10 per million input tokens, but prompts over 100K cost 5x. Compared with DeepSeek V4.1 Flash, GLM-5.3-Flash and MiMo-V2.6-Flash (10/08).

Data checked on October 8, 2026 (called "as of 10/08" below). Prices come from each vendor's official pricing page. The Intelligence Index and benchmark scores come from the model and comparison pages of Artificial Analysis (AA). BazaarLink prices come from the public model catalog (/api/v1/models). Vendors can change prices or update scores at any time, so check the source pages before you rely on a number.

Claude Haiku 5.5 launched on October 7 with input at US$0.10 per million tokens. It is not the only cheap small model: DeepSeek V4.1 Flash, GLM-5.3-Flash and MiMo-V2.6-Flash sit in the same price band. This article puts price and AA-measured capability for all four in one place, and explains why a low unit price does not mean a low total bill.

The short version

  • Haiku 5.5 has the lowest input price, but only for prompts of 100K tokens or less. Above 100K tokens, both input and output prices become 5x (US$0.50 / US$2.50).
  • AA Intelligence Index: Haiku 5.5 scores 43, the highest of the four, but only 1 point ahead of GLM-5.3-Flash (42). That is Haiku 5.5 at its highest reasoning effort (Max). At Medium it scores 34, below the other three.
  • A low unit price does not mean a cheap task. Haiku 5.5 (Max) writes about 162k output tokens per task, the most of the four. AA measured an average cost of US$0.21 per task, while MiMo-V2.6-Flash costs only US$0.06.
  • Terminal and coding agents (Terminal-Bench 4.0): Haiku 5.5 and GLM-5.3-Flash tie for the top (33%). SaaS workflows (AutomationBench-AA): Haiku 5.5 is lowest (35%), but AA says that number is probably too low (details below).

Price comparison (official list prices)

Unit: US dollars per million tokens. As of 10/08. The three columns "Official input / output / cache hit" are each vendor's official list price; the two right-hand columns are BazaarLink prices.

ModelOfficial inputOfficial outputOfficial cache hitBazaarLink inputBazaarLink output
Claude Haiku 5.5 (prompt ≤100K)0.100.500.010.100.50
Claude Haiku 5.5 (prompt >100K)0.502.500.050.502.50
DeepSeek V4.1 Flash (peak)0.301.200.0060.301.20
DeepSeek V4.1 Flash (off-peak)0.150.600.003——
GLM-5.3-Flash0.150.500.030.150.50
MiMo-V2.6-Flash0.140.280.00280.140.28

"—" means BazaarLink does not have that price.

A few things to watch:

  • Haiku 5.5 has two price tiers based on prompt length. Anthropic's pricing page says prompts over 100,000 tokens use a higher set of prices, and the BazaarLink catalog also lists a tier price above 100,000 tokens. Long documents, long conversations and large tool outputs cross that line easily, so estimate cost with your real prompt length.
  • DeepSeek V4.1 Flash has peak and off-peak prices. DeepSeek's pricing page says off-peak prices are half of peak prices. Peak hours are Monday to Friday (excluding Chinese public holidays), 01:00–04:00 and 06:00–10:00 UTC. The BazaarLink catalog lists DeepSeek V4.1 Flash at a single price, equal to the official peak price.
  • GLM-5.3-Flash cache storage is listed as "limited-time free" on Z.ai's pricing page and may be charged later.
  • MiMo-V2.6-Flash: Xiaomi's official model page lists both CNY and USD prices: input ¥1, output ¥2, cache hit ¥0.02 per million tokens. The USD prices are the ones in the table.
  • Context length: in the BazaarLink catalog all four models have about 1 million tokens (1,048,576 for MiMo-V2.6-Flash).

Haiku 5.5 also has a 5-minute cache write at US$0.125 and a 1-hour cache write at US$0.20 (prompt ≤100K, per million tokens). The other three vendors' official pages do not list the same write tiers, so the table leaves cache writes out.

AA Intelligence Index and benchmark breakdown

The AA Intelligence Index (v4.3.2) combines 10 evaluations. On AA's pages, Haiku 5.5 and DeepSeek V4.1 Flash are labeled Max reasoning effort; GLM-5.3-Flash and MiMo-V2.6-Flash show no reasoning-effort label. As of 10/08:

ModelAA Intelligence IndexBar (each block ≈ 2 points)
Reference: Claude Sonnet 5.556████████████████████████████
Claude Haiku 5.5 (Max)43█████████████████████▌
GLM-5.3-Flash42█████████████████████
DeepSeek V4.1 Flash (Max)39███████████████████▌
MiMo-V2.6-Flash38███████████████████
Reference: Claude Haiku 4.5 (previous generation)17████████▌

Haiku 5.5 is 26 points above the previous generation, Haiku 4.5. A 1-point gap between Haiku 5.5 and GLM-5.3-Flash should not be read as a clear win.

Benchmark breakdown (measured by AA; fewer output tokens per task is cheaper, and cost per task is AA's average for running the Intelligence Index):

ItemHaiku 5.5 (Max)GLM-5.3-FlashDeepSeek V4.1 Flash (Max)MiMo-V2.6-Flash
Terminal-Bench 4.033%33%27%23%
AA-Omniscience Index117−5−13
AutomationBench-AA35%*60%69%64%
Average output tokens per task162k69k89k78k
Average cost per taskUS$0.21US$0.25US$0.27US$0.06
  • Terminal-Bench 4.0 (terminal and coding agents, higher is better): AA describes it as 66 tasks of complex terminal work across software, machine learning, science and other areas.
  • AA-Omniscience (higher means more reliable): the index ranges from −100 to 100. AA says it "rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer"; 0 means as many correct answers as incorrect ones, and negative scores mean more incorrect than correct.
  • AutomationBench-AA (SaaS workflows, higher is better): it measures how often an agent completes tasks in simulated SaaS application environments. A task that breaks a guardrail scores zero.

* AA's note: Haiku 5.5's AutomationBench-AA score of 35% is likely understated. AA says that during pre-release testing a safety refusal issue made the model over-refuse. Anthropic is working on a fix, and AA plans to re-run the evaluation once it lands and expects the score to rise (this is AA's statement on X, as reported by OfficeChai). Until the re-run, this cell is not suitable for comparing with the other three.

Unit price is not total price: cost per task

AA runs the same Intelligence Index tasks and reports each model's average output volume and cost per task:

ModelOutput price (per million)Output tokens per taskAverage cost per task
Claude Haiku 5.5 (Max)US$0.50162kUS$0.21
DeepSeek V4.1 Flash (Max)US$1.20 (peak)89kUS$0.27
GLM-5.3-FlashUS$0.5069kUS$0.25
MiMo-V2.6-FlashUS$0.2878kUS$0.06

Haiku 5.5's output price equals GLM-5.3-Flash's, but at Max effort it writes about 2.3 times as many output tokens per task, which eats part of the unit-price advantage. MiMo-V2.6-Flash has both a low unit price and low output volume, so it has the lowest cost per task. Your real cost will vary with peak versus off-peak pricing, cache hit rate and prompt length.

Haiku 5.5's reasoning effort changes both score and cost

Haiku 5.5 lets you choose a reasoning effort, and the BazaarLink catalog lists the default as medium. AA scores each effort separately:

Reasoning effortAA Intelligence IndexAverage cost per task
Max43US$0.21
High38US$0.08
Medium34US$0.05
Low29US$0.02

So "Haiku 5.5 scores 43" means Max. At the default medium, AA's score is 34, below GLM-5.3-Flash (42), DeepSeek V4.1 Flash (39) and MiMo-V2.6-Flash (38), although the cost per task is also much lower. To compare the four models, fix the reasoning effort and test with your own tasks.

How to choose

  • Prompts mostly under 100K tokens, and you want the highest AA score of the four: Haiku 5.5 (Max). First check that your task can accept its higher output volume per task.
  • Cost per task comes first: by AA's cost per task, MiMo-V2.6-Flash is lowest; its AutomationBench-AA score is 64% and its Terminal-Bench 4.0 score is 23%.
  • Workflow-style SaaS automation: on AA's AutomationBench-AA today, DeepSeek V4.1 Flash is highest (69%), then MiMo-V2.6-Flash (64%) and GLM-5.3-Flash (60%). Do not use Haiku 5.5's number until AA re-runs it.
  • Mostly long prompts (over 100K tokens): Haiku 5.5's price is 5x (input US$0.50, output US$2.50) and is no longer the lowest of the four. The other three have no such tier.

All of this is based on one AA benchmark mix and official list prices, not a guarantee for your task. Pick a set of representative tasks, fix the reasoning effort, run them side by side, and then decide.

All four models can be called through the same OpenAI-compatible endpoint. Just change model:

for MODEL in claude-haiku-5.5 deepseek-v4.1-flash glm-5.3-flash mimo-v2.6-flash; do
  curl -s https://bazaarlink.ai/api/v1/chat/completions \
    -H "Authorization: Bearer $BAZAARLINK_API_KEY" \
    -H "Content-Type: application/json" \
    -d "{\"model\": \"$MODEL\", \"messages\": [{\"role\": \"user\", \"content\": \"Explain prompt caching in three sentences.\"}]}"
done

For live prices and specs, see the BazaarLink model catalog. If you are weighing an upgrade from Haiku 4.5, read Should you move from Claude Haiku 4.5 to 5.5?

FAQ

Which setting is Haiku 5.5's AA Intelligence Index of 43 measured at? Max reasoning effort. AA scores each effort separately: Max 43, High 38, Medium 34, Low 29.

Why is Haiku 5.5's AutomationBench-AA score only 35%? AA says the score is likely understated, because a safety refusal issue in pre-release testing made the model over-refuse. Anthropic is working on a fix, and AA plans to re-run the evaluation after it lands.

Does DeepSeek V4.1 Flash have peak and off-peak prices on BazaarLink? The BazaarLink catalog lists a single price: US$0.30 input and US$1.20 output, equal to the official peak price. On DeepSeek's own API the off-peak price is half the peak price.

What happens when a prompt goes over 100K tokens? For Haiku 5.5, prompts over 100,000 tokens pay 5x on both input and output (US$0.50 / US$2.50). The other three models have no such tier.

Do these scores represent my task? Not directly. AA's index is its own benchmark mix, so test side by side with your own data.


Sources: Anthropic pricing, DeepSeek pricing, Z.ai pricing, Xiaomi MiMo-V2.6-Flash model page, AA comparison pages Haiku 5.5 vs GLM-5.3-Flash, vs DeepSeek V4.1 Flash and vs MiMo-V2.6-Flash, and AA's Haiku 5.5 (Max), High, Medium, Low and Haiku 5.5 release page. Data checked on October 8, 2026.

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
Claude · Haiku 5.5 · Haiku 4.5 · Model Comparison · Pricing
Haiku 4.5 or 5.5? $0.10/$0.50 Specs and NT$ Cost Estimates
Gemini 4 Argon · Vending-Bench 2 · AI agent · Andon Labs · agent safety
Gemini 4 Argon: #3 on Vending-Bench 2, accused of lying
Claude · Sonnet 5.5 · Sonnet 5 · Anthropic · API migration
Sonnet 5.5: 30% Faster and Cheaper? API Changes Before Upgrading
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.