NVIDIA
Available in Shortly AI
Nemotron 3 Ultra
Released June 4, 2026
ReasoningTool Use
A 550B parameter (55B active) open reasoning model from NVIDIA, built for long-running agent workflows. It uses a hybrid Mamba-Transformer MoE architecture and supports a 1M token context window.
- Context window
- 1M tokens
- Maximum output
- 65K tokens
- Input pricing
- $0.60/M
- Output pricing
- $2.40/M
Independent benchmark
LiveBench historical snapshot
51.8overall
Historical benchmark data from LiveBench, release 2026-01-08. Scores are pinned and are not fetched on page load.
37.5
71.3
46.7
54.5
42.0
52.2
58.2
| Category | Score |
|---|---|
| Reasoning | 37.543 |
| Coding | 71.341 |
| Agentic Coding | 46.667 |
| Mathematics | 54.518 |
| Data Analysis | 42.014 |
| Language | 52.180 |
| Instruction Following | 58.192 |
Benchmarks measure selected evaluation tasks and do not guarantee performance for every prompt or workflow.
FAQ
Similar AI models
Inkling SmallThinkingmachinesInkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency.Gemini 3.5 Flash LiteGoogleGemini 3.5 Flash Lite features upgraded agentic capabilities, making the model ideal for subagents in complex workflows.Gemini 3.6 FlashGoogleGemini 3.6 Flash delivers higher quality across coding, agentic workflows, and web development with reduced token consumption and fewer model calls compared to previous model iterations.
Available with one subscription
Start with Nemotron 3 Ultra in Shortly AI
Keep your conversations, files and model choices together in one workspace.