Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4

NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 is NVIDIA's 550B-total, 55B-active LatentMoE reasoning model. According to the model card, it combines Mamba-2, MoE, and Attention with MTP and NVFP4 efficiency for agentic and long-context work.

Publisher
nvidia
Task
text-generation
Model type
nemotron_h
License
openmdw-1.1
Library
transformers
Publication status
Accepted · not indexed

Model overview

NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 is published by nvidia as a nemotron_h text-generation model. The captured configuration identifies NemotronHForCausalLM. According to the model card, it is a LatentMoE hybrid of Mamba-2, MoE, and Attention with Multi-Token Prediction, described as a frontier-scale reasoning and chat model with 550B total and 55B active parameters.

Recorded capabilities

550B LatentMoE with 55B active

According to the model card, the model has 550B total and 55B active parameters in a Mamba-2, MoE, and Attention hybrid with Multi-Token Prediction layers.

Up-to-1M context with reasoning control

According to the model card, maximum context reaches up to 1M tokens, with configurable reasoning via enable_thinking and medium-effort or budgeted modes.

Card-reported agentic and reasoning evals

According to the card's Nemo Evaluator table, the NVFP4 variant reports SWE-Bench Verified 69.5, GPQA 87.9, IFBench prompt 82.3, and RULER 1M 94.0.

Four-stage post-training recipe

According to the model card, training spans approximately 20T-token pretraining, supervised fine-tuning, asynchronous GRPO reinforcement learning with RLHF, and multi-domain on-policy distillation.

Nemotron-H Transformers checkpoint

Captured config identifies NemotronHForCausalLM and nemotron_h, with Safetensors metadata reporting 302826566168 parameters and transformers support plus latent-moe and mtp tags.

OpenMDW-1.1 terms

According to the model card, use is governed by the OpenMDW-1.1 license, with card data recording license other.

Use cases in the source record

  • Configurable reasoning chat with thinking on, off, or medium-effort modes, plus tool-calling workflows using the card's OpenAI-compatible API examples.
  • Long-context analysis and high-stakes RAG over large documents and codebases, deployed on the card's stated 4x B200 or multi-node GB200/GB300 minimum hardware.

Limitations and unknowns

  • Provider state is historical snapshot data with mixed live and error statuses, not independently refreshed current availability.
  • Benchmark figures are publisher-reported via Nemo Evaluator harnesses and were not independently verified by Ethen.
  • Captured Safetensors size and the card's 550B total claim describe different checkpoint representations and should not be conflated without validation.

Source and provenance

Source: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4

Captured: Unknown. Processed: 2026-09-07T19:34:54.773359+00:00.

NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 Model Summary Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, and Chinese Best For Frontier reasoning, complex agentic workflows, long-context analysis, tool use, multilingual reasoning, high-stakes RAG Reasoning Mode Configurable on/off via chat template ( enable_thinking=True/False ) License OpenMDW License Ag…

F001F002F003F004F005F006F007F008F009F010F011F012F016F017F018F019F020F021F023F024F025F026F029F032F035F036F037