NVIDIA
Available in Shortly AI
Nvidia Nemotron Nano 12B V2 VL
Released October 28, 2025
ReasoningVision (Image)Tool Use
The model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities.
- Context window
- 131.1K tokens
- Maximum output
- 131.1K tokens
- Input pricing
- $0.20/M
- Output pricing
- $0.60/M
FAQ
Similar AI models
Qwen 3.7 FlashAlibabaThe Qwen3.7 native vision-language Flash model series delivers a comprehensive upgrade over 3.6-Flash in multimodal understanding and agent execution. This model particularly excels in enhanced multimodal foundations with stronger universal object recognition, further improved real-world perception and spatial intelligence, significantly upgraded multimodal agent capabilities for Search Agent and CI Agent scenarios with more stable end-to-end task execution, as well as optimized multimodal coding for a smoother vibe coding experience.GPT 5.6 LunaOpenAIGPT-5.6 Luna is a fast, affordable GPT-5.6 model that brings strong capability at the lowest cost in the series.Hy3TencentTencent Hy3 is an open-source Mixture-of-Experts (MoE) large language model developed by Tencent's Hunyuan team. It has 295B total parameters, 21B active parameters, a 3.8B-parameter MTP layer, and a 256K context window. Hy3 is designed for agentic workflows, coding, reasoning, long-context understanding, tool-use scenarios, document analysis, and structured generation.
The model improves on Hy3 Preview with stronger tool-call reliability, better output-format stability, improved long-context retention, and lower hallucination rates in Tencent's internal evaluations.
Available with one subscription
Start with Nvidia Nemotron Nano 12B V2 VL in Shortly AI
Keep your conversations, files and model choices together in one workspace.