MoE Mamba hybrid
According to the model card, the hybrid MoE architecture combines Mamba-2 and MoE layers with select attention layers, totaling 30B with 3B active.
Open Source Model Profile · nvidia
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 is a 31.58B-parameter Nemotron hybrid text-generation model from nvidia. According to the model card, it is the BF16 reference with 3B active parameters for customization.
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 is published by nvidia as a text-generation model. The captured configuration identifies NemotronHForCausalLM, and Safetensors metadata reports about 31.58B parameters. According to the model card, it is the full-precision BF16 reference for the 30B-total, 3B-active Lightning hybrid, intended mainly for customization.
According to the model card, the hybrid MoE architecture combines Mamba-2 and MoE layers with select attention layers, totaling 30B with 3B active.
According to the model card, the model supports up to 1M tokens, with validated recipes for single H100, multi-GPU, GB200, and B200 setups.
According to the model card, these BF16 weights are the starting point for post-training, distillation, domain adaptation, and quantized variants.
According to the model card, reasoning mode is configurable through the chat template and the card documents function-call tool examples.
According to the model card, the publisher gives vLLM and SGLang commands, hardware pairings, and speculative-decoding options including DSpark.
Source: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Captured: Unknown. Processed: 2026-09-07T19:34:54.865374+00:00.
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Model Summary Total Parameters 30B (3B active) Architecture MoE — Mamba-2 + MoE + Attention hybrid Precision BF16 (full-precision reference weights) Context Length Up to 1M tokens (for single H100 deployment, we use 256K) Single-GPU Deployment 1× H100 80GB (or 1× A100 80GB) Supported Hardware NVIDIA Blackwell (GB200, B200); NVIDIA Hopper (H100, H200); NVIDIA Ampere (A100) Supported Languages English (and coding languages), Spanish, French, German, Italian, Japanese Speculative Decoding DSpark for Low Concurrency Data Centre Deployments — Read more below Reasoning Mode Configurable on/off vi…
F001F002F003F004F005F006F007F009F010F011F016F017F020F021F022F028F029F031F032F033F035F036