Cogito V2 Preview Llama 109b Moe
An instruction-tuned, hybrid-reasoning Mixture-of-Experts model built on Llama-4-Scout-17B-16E. Cogito v2 can answer directly or engage an extended “thinking” phase, with alignment guided by Iterated Distillation & Amplification (IDA). It targets coding, STEM, instruction following, and general helpfulness, with stronger multilingual, tool-calling, and reasoning performance than size-equivalent baselines. The model supports long-context use (up to 10M tokens) and standard Transformers workflows. Users can control the reasoning behaviour with the `reasoning` `enabled` boolean. [Learn more in our docs](https:///docs/use-cases/reasoning-tokens#enable-reasoning-with-default-config)
Pricing
💱 匯率 USD/NTD = 32.34 · 不含稅 · 更新於 2026/07/22
Technical specifications
Quick Start
Use Cogito V2 Preview Llama 109b Moe via BazaarLink API — just change the base_url:
from openai import OpenAI
client = OpenAI(
base_url="https://bazaarlink.ai/api/v1",
api_key="sk-bl-YOUR_API_KEY",
)
response = client.chat.completions.create(
model="cogito-v2-preview-llama-109b-moe",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://bazaarlink.ai/api/v1",
apiKey: "sk-bl-YOUR_API_KEY",
});
const response = await client.chat.completions.create({
model: "cogito-v2-preview-llama-109b-moe",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);Why use Cogito V2 Preview Llama 109b Moe via BazaarLink?
- ✓USD billing (TWD-quoted) + unified invoices — no foreign credit card needed for Taiwan teams
- ✓OpenAI-compatible API — zero code changes required
- ✓Automatic failover — multi-provider redundancy for the same model
- ✓Chinese-language support — local team, instant help
Frequently Asked Questions
What is Cogito V2 Preview Llama 109b Moe?
An instruction-tuned, hybrid-reasoning Mixture-of-Experts model built on Llama-4-Scout-17B-16E. Cogito v2 can answer directly or engage an extended “thinking” phase, with alignment guided by Iterated Distillation & Amplification (IDA). It targets coding, STEM, instruction following, and general helpfulness, with stronger multilingual, tool-calling, and reasoning performance than size-equivalent baselines. The model supports long-context use (up to 10M tokens) and standard Transformers workflows. Users can control the reasoning behaviour with the `reasoning` `enabled` boolean. [Learn more in our docs](https:///docs/use-cases/reasoning-tokens#enable-reasoning-with-default-config)
How much does the Cogito V2 Preview Llama 109b Moe API cost?
Cogito V2 Preview Llama 109b Moe pricing is available on this page. BazaarLink bills in TWD with no additional markup over the provider's list price.
How do I use Cogito V2 Preview Llama 109b Moe with the OpenAI SDK?
Set base_url to "https://bazaarlink.ai/api/v1" and use model ID "cogito-v2-preview-llama-109b-moe". All OpenAI SDK methods (chat.completions, embeddings, streaming) work without code changes.
What is the context window for Cogito V2 Preview Llama 109b Moe?
Cogito V2 Preview Llama 109b Moe supports a context window of 131,072 tokens.
Is Cogito V2 Preview Llama 109b Moe available for free?
Cogito V2 Preview Llama 109b Moe is a paid model. BazaarLink offers free trial credits on registration so you can test it without a credit card.