BazaarLinkBazaarLink
Sign in

Deepseek R1 Distill Llama 8b — Retired model

Retired modelThis model has reached end of life and is no longer available through BazaarLink. Historical specifications and independent benchmarks remain below. Retired on May 1, 2026.

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including: - AIME 2024 pass@1: 50.4 - MATH-500 pass@1: 89.1 - CodeForces Rating: 1205 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models. Hugging Face: - [Llama-3.1-8B](https://huggingface.co/meta-llama/Llama-3.1-8B) - [DeepSeek-R1-Distill-Llama-8B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B) |

ProviderDeepSeek
ReleasedFebruary 2025
Model IDdeepseek/deepseek-r1-distill-llama-8b

Technical specifications

Context window
— tokens
Reasoning
Input
Text
Output
Text

Independent benchmarks

Scores below come from Artificial Analysis, an independent third party. BazaarLink does not participate in the testing.

Intelligence 7
Coding
Math 41
MMLU 54
GPQA 30
This modelCategory leader
Overall intelligence
#511out of 638 models
Aggregated across academic benchmarks
Intelligence
7
Mid-tier
This model
7
Category leader
53
Median
12
Math
41
This model
41
Category leader
99
Median
53
MMLU Pro
54%
This model
54%
Category leader
90%
Median
75%
GPQA
30%
This model
30%
Category leader
94%
Median
69%
LiveCodeBench
23%
This model
23%
Category leader
42%
Median
42%
HLE
4%
This model
4%
Category leader
53%
Median
7%
SciCode
12%
This model
12%
Category leader
60%
Median
34%
IFBench
18%
This model
18%
Category leader
83%
Median
44%
AA-LCR
0%
This model
0%
Category leader
76%
Median
40%

Top 5 — CodingCoding Index leaderboard

2Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)?/claude-fable-5-1-xhigh81$10 / $50
3Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)?/claude-fable-5-1-high79$10 / $50
4GPT-5.6 Sol (xhigh)openai/gpt-5-6-sol-xhigh78$4 / $20
AASource: Artificial Analysis · Independent third-party evaluation; not influenced by BazaarLinkView full evaluation on AA →
← All Models

Frequently Asked Questions

Can I still use the retired Deepseek R1 Distill Llama 8b model?

No. Deepseek R1 Distill Llama 8b has been retired and is retained here only as a historical reference. See the active alternatives linked on this page.

Relay integrity checks

No integrity-check data is available for this model yet.

BazaarLink Probe verifies endpoints that claim to serve this model and flags behavior that differs from the reference baseline.

View the full integrity report →

Related models and links

More from this provider: DeepSeek

Deepseek V3.2Deepseek V4 FlashDeepseek V4 ProDeepseek Chat V3 0324Deepseek V4 Pro 0813Deepseek V3.2 Exp
ModelsAPI documentationAPI latency probe
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.