Toggle navigation

NVIDIA

Available in Shortly AI

Nemotron 3.5 Lightning 30B

Released August 11, 2026

ReasoningTool Use

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.

Context window
262.1K tokens
Maximum output
131.1K tokens
Input pricing
$0.05/M
Output pricing
$0.20/M