Skip to content

EthenEthenEthen

Open Source Model Profile · nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 is a 31.58B-parameter Nemotron hybrid text-generation model from nvidia. According to the model card, it is the BF16 reference with 3B active parameters for customization.

Publisher
nvidia
Task
text-generation
Model type
nemotron_h
License
openmdw-1.1
Library
transformers
Publication status
Approved for indexing

Model overview

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 is published by nvidia as a text-generation model. The captured configuration identifies NemotronHForCausalLM, and Safetensors metadata reports about 31.58B parameters. According to the model card, it is the full-precision BF16 reference for the 30B-total, 3B-active Lightning hybrid, intended mainly for customization.

Recorded capabilities

MoE Mamba hybrid

According to the model card, the hybrid MoE architecture combines Mamba-2 and MoE layers with select attention layers, totaling 30B with 3B active.

1M-token context design

According to the model card, the model supports up to 1M tokens, with validated recipes for single H100, multi-GPU, GB200, and B200 setups.

Customization reference weights

According to the model card, these BF16 weights are the starting point for post-training, distillation, domain adaptation, and quantized variants.

Reasoning and tool use

According to the model card, reasoning mode is configurable through the chat template and the card documents function-call tool examples.

Documented vLLM and SGLang paths

According to the model card, the publisher gives vLLM and SGLang commands, hardware pairings, and speculative-decoding options including DSpark.

Use cases in the source record

  • Post-training and customization work such as SFT, RL, distillation, domain adaptation, and quantized-variant construction.
  • Full-precision research and evaluation using the documented vLLM and SGLang serving recipes.

Limitations and unknowns

  • Evaluation coverage is described only as a release suite with NeMo Gym and NeMo Evaluator recipes; no extracted numeric scores are stated here.
  • According to the model card, BF16 on a single H100 is memory-bound to about 256K, so full 1M use needs larger multi-GPU or newer-hardware setups.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Captured: Unknown. Processed: 2026-09-07T19:34:54.865374+00:00.

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Model Summary Total Parameters 30B (3B active) Architecture MoE — Mamba-2 + MoE + Attention hybrid Precision BF16 (full-precision reference weights) Context Length Up to 1M tokens (for single H100 deployment, we use 256K) Single-GPU Deployment 1× H100 80GB (or 1× A100 80GB) Supported Hardware NVIDIA Blackwell (GB200, B200); NVIDIA Hopper (H100, H200); NVIDIA Ampere (A100) Supported Languages English (and coding languages), Spanish, French, German, Italian, Japanese Speculative Decoding DSpark for Low Concurrency Data Centre Deployments — Read more below Reasoning Mode Configurable on/off vi…

F001F002F003F004F005F006F007F009F010F011F016F017F020F021F022F028F029F031F032F033F035F036