Meta
Available in Shortly AI
Llama 4 Scout 17B Instruct
Released April 5, 2025
Vision (Image)Tool Use
Llama 4 Scout is the best multimodal model in the world in its class and is more powerful than our Llama 3 models, while fitting in a single H100 GPU. Additionally, Llama 4 Scout supports an industry-leading context window of up to 10M tokens.
- Context window
- 128K tokens
- Maximum output
- 8.2K tokens
- Input pricing
- $0.17/M
- Output pricing
- $0.66/M
FAQ
Similar AI models
GLM 5.3 FlashZ.aiGLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.DeepSeek V4 Flash Vision ExpDeepSeekDeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks.Nemotron 3.5 Lightning 30BNVIDIANVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA.
The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
Available with one subscription
Start with Llama 4 Scout 17B Instruct in Shortly AI
Keep your conversations, files and model choices together in one workspace.