Longcat Flash Chat
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce communication overhead and achieve high throughput while maintaining training stability through advanced scaling strategies such as hyperparameter transfer, deterministic computation, and multi-stage optimization. This release, LongCat-Flash-Chat, is a non-thinking foundation model optimized for conversational and agentic tasks. It supports long context windows up to 128K tokens and shows competitive performance across reasoning, coding, instruction following, and domain benchmarks, with particular strengths in tool use and complex multi-step interactions.
Pricing
💱 匯率 USD/NTD = 32.41 · 不含稅 · 更新於 2026/07/23
Technical specifications
Quick Start
Use Longcat Flash Chat via BazaarLink API — just change the base_url:
from openai import OpenAI
client = OpenAI(
base_url="https://bazaarlink.ai/api/v1",
api_key="sk-bl-YOUR_API_KEY",
)
response = client.chat.completions.create(
model="longcat-flash-chat",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://bazaarlink.ai/api/v1",
apiKey: "sk-bl-YOUR_API_KEY",
});
const response = await client.chat.completions.create({
model: "longcat-flash-chat",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);Why use Longcat Flash Chat via BazaarLink?
- ✓USD billing (TWD-quoted) + unified invoices — no foreign credit card needed for Taiwan teams
- ✓OpenAI-compatible API — zero code changes required
- ✓Automatic failover — multi-provider redundancy for the same model
- ✓Chinese-language support — local team, instant help
Frequently Asked Questions
What is Longcat Flash Chat?
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce communication overhead and achieve high throughput while maintaining training stability through advanced scaling strategies such as hyperparameter transfer, deterministic computation, and multi-stage optimization. This release, LongCat-Flash-Chat, is a non-thinking foundation model optimized for conversational and agentic tasks. It supports long context windows up to 128K tokens and shows competitive performance across reasoning, coding, instruction following, and domain benchmarks, with particular strengths in tool use and complex multi-step interactions.
How much does the Longcat Flash Chat API cost?
Longcat Flash Chat costs $0.2000 per 1K input tokens and $0.8000 per 1K output tokens when accessed through BazaarLink.
How do I use Longcat Flash Chat with the OpenAI SDK?
Set base_url to "https://bazaarlink.ai/api/v1" and use model ID "longcat-flash-chat". All OpenAI SDK methods (chat.completions, embeddings, streaming) work without code changes.
What is the context window for Longcat Flash Chat?
Longcat Flash Chat supports a context window of 131,072 tokens.
Is Longcat Flash Chat available for free?
Longcat Flash Chat is a paid model. BazaarLink offers free trial credits on registration so you can test it without a credit card.