BazaarLinkBazaarLink
Sign in

Llama 3.3 Nemotron Super 49b V1 — Retired model

Retired modelThis model has reached end of life and is no longer available through BazaarLink. Historical specifications and independent benchmarks remain below. Retired on May 1, 2026.

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) optimized for advanced reasoning, conversational interactions, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta's Llama-3.3-70B-Instruct, it employs a Neural Architecture Search (NAS) approach, significantly enhancing efficiency and reducing memory requirements. This allows the model to support a context length of up to 128K tokens and fit efficiently on single high-performance GPUs, such as NVIDIA H200. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see [Usage Recommendations](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1#quick-start-and-usage-recommendations) for more.

ProviderNVIDIA
ReleasedApril 2025
Model IDnvidia/llama-3.3-nemotron-super-49b-v1

Technical specifications

Context window
131K tokens
Reasoning
Input
Text
Output
Text
← All Models

Frequently Asked Questions

Can I still use the retired Llama 3.3 Nemotron Super 49b V1 model?

No. Llama 3.3 Nemotron Super 49b V1 has been retired and is retained here only as a historical reference. See the active alternatives linked on this page.

Relay integrity checks

No integrity-check data is available for this model yet.

BazaarLink Probe verifies endpoints that claim to serve this model and flags behavior that differs from the reference baseline.

View the full integrity report →

Related models and links

More from this provider: NVIDIA

Nemotron 3 Super 120b A12bNemotron 3 Nano 30b A3bGlm 4.7Glm 5Glm 5.2Deepseek V3.2
ModelsAPI documentationAPI latency probe
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.