BazaarLinkBazaarLink
Sign in
T

Llama 3.1 Swallow 8b Instruct V0.3 — Retired model

Retired modelThis model has reached end of life and is no longer available through BazaarLink. Historical specifications and independent benchmarks remain below. Retired on May 1, 2026.

Llama 3.1 Swallow 8B is a large language model that was built by continual pre-training on the Meta Llama 3.1 8B. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. Swallow used approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre-training. The instruction-tuned models (Instruct) were built by supervised fine-tuning (SFT) on the synthetic data specially built for Japanese.

ProviderTtokyotech-llm
ReleasedApril 2025
Model IDtokyotech-llm/llama-3.1-swallow-8b-instruct-v0.3

Technical specifications

Context window
16K tokens
Reasoning
Input
Text
Output
Text
← All Models

Frequently Asked Questions

Can I still use the retired Llama 3.1 Swallow 8b Instruct V0.3 model?

No. Llama 3.1 Swallow 8b Instruct V0.3 has been retired and is retained here only as a historical reference. See the active alternatives linked on this page.

Relay integrity checks

No integrity-check data is available for this model yet.

BazaarLink Probe verifies endpoints that claim to serve this model and flags behavior that differs from the reference baseline.

View the full integrity report →

Related models and links

Related models and links

Glm 4.7Glm 5Glm 5.2Deepseek V3.2Deepseek V4 ProGlm 5.1
ModelsAPI documentationAPI latency probe
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.