NVIDIA
Available in Shortly AI
Nemotron 3.5 Lightning 30B
Released August 11, 2026
ReasoningTool Use
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
- Context window
- 262.1K tokens
- Maximum output
- 131.1K tokens
- Input pricing
- $0.05/M
- Output pricing
- $0.20/M
FAQ
Similar AI models
GLM 5.3 FlashZ.aiGLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.DeepSeek V4 Flash Vision ExpDeepSeekDeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks.Ling 3.0 FlashInclusionaiLing 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Available with one subscription
Start with Nemotron 3.5 Lightning 30B in Shortly AI
Keep your conversations, files and model choices together in one workspace.