Nemotron 3 Ultra
NVIDIA’s largest open model: hybrid Mamba-Transformer MoE, 55B active.
Nemotron 3 Ultra is NVIDIA’s open frontier-reasoning and orchestration model, with 55B active out of 550B total parameters, built on a hybrid Transformer-Mamba mixture-of-experts architecture.
- Provider
- NVIDIA
- Type
- Language model
- Released
- Jun 4, 2026
- Context window
- 262K tokens
- Max output
- 16K tokens
- Price
- $0.50 input / $2.20 output per 1M tokens
- Input
- text
- Output
- text
- Open weights
- Yes
- License
- NVIDIA Open Model License
- Knowledge cutoff
- Sep 30, 2025
Best for
- Reasoning
- Agents
- Enterprise
- Local & on-device
Strengths
- Open weights
- Hybrid Mamba for long-context throughput
- Orchestration
More from NVIDIA
- Nemotron 3 Super — NVIDIA. Open hybrid Mamba-Transformer MoE for multi-agent applications.
- Nemotron 3.5 Lightning — NVIDIA. Tiny open 30B-A3B model with the lowest first-token latency in the index.
- Nemotron 3 Embed 1B — NVIDIA. Small open embedding model for high-throughput enterprise retrieval.
- Cosmos 3 — NVIDIA. NVIDIA’s open omni-modal world foundation model for physical AI.
- Isaac GR00T N1.7 — NVIDIA. Open 3B cross-embodiment VLA for humanoids and manipulators.
- Nemotron 3.5 Content Safety — NVIDIA. Compact 4B multimodal guardrail model for LLM and VLM traffic.
Tools that use it
- NVIDIA NIM — NVIDIA. Prebuilt, optimised inference microservices for NVIDIA GPUs.