BazaarLinkBazaarLink
Sign in
M

Longcat Flash Chat — Retired model

Retired modelThis model has reached end of life and is no longer available through BazaarLink. Historical specifications and independent benchmarks remain below. Retired on May 1, 2026.

LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce communication overhead and achieve high throughput while maintaining training stability through advanced scaling strategies such as hyperparameter transfer, deterministic computation, and multi-stage optimization. This release, LongCat-Flash-Chat, is a non-thinking foundation model optimized for conversational and agentic tasks. It supports long context windows up to 128K tokens and shows competitive performance across reasoning, coding, instruction following, and domain benchmarks, with particular strengths in tool use and complex multi-step interactions.

ProviderMmeituan
ReleasedSeptember 2025
Model IDmeituan/longcat-flash-chat

Technical specifications

Context window
131K tokens
Reasoning
Input
Text
Output
Text
← All Models

Frequently Asked Questions

Can I still use the retired Longcat Flash Chat model?

No. Longcat Flash Chat has been retired and is retained here only as a historical reference. See the active alternatives linked on this page.

Relay integrity checks

No integrity-check data is available for this model yet.

BazaarLink Probe verifies endpoints that claim to serve this model and flags behavior that differs from the reference baseline.

View the full integrity report →

Related models and links

Related models and links

Glm 4.7Glm 5Glm 5.2Deepseek V3.2Deepseek V4 ProGlm 5.1
ModelsAPI documentationAPI latency probe
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.