NVIDIA Nemotron 3 Ultra technical report
NVIDIA released a 550B-total, 55B-active hybrid Mamba-attention MoE model with an open post-training stack aimed at agentic workloads
Proud to see Nemotron 3 Ultra out! From Nano → Super → Ultra, we’ve kept pushing the post-training stack: SFT → multi-env RLVR → MOPD with many specialized teachers. The result: fast, token-efficient, open, and strong on agentic tasks. T