Model directory
Find the right AI model
Explore leading AI models, compare their strengths, and find the right one for your next task.
Top model per creator
The highest-scoring currently available model from each creator, ordered by model release date. Scores are comparable within this release only.
78.5Release order 179.2Release order 271.9Release order 375.3Release order 481.1Release order 583.0Release order 673.2Release order 774.2Release order 877.0Release order 9
| Release order | Model | Overall score |
|---|---|---|
| 1 | Qwen 3.8 Max | 78.461 |
| 2 | Kimi K3 | 79.193 |
| 3 | Inkling | 71.923 |
| 4 | Muse Spark 1.1 | 75.300 |
| 5 | GPT 5.6 Sol | 81.054 |
| 6 | Claude Fable 5 | 82.971 |
| 7 | GLM 5.2 | 73.157 |
| 8 | DeepSeek V4 Flash 0731 | 74.171 |
| 9 | Gemini 3.1 Pro Preview | 76.952 |
Models
117 modelsCapabilities
Qwen 3.8 MaxAlibabaQwen 3.8 Max is a 2.4-trillion-parameter MoE model delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days. Handles hundreds of specialized tasks across legal, financial, design, and other professional domains, producing production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. In long-horizon tasks, plans autonomously, iterates through closed feedback loops, and continuously evolves.2026-08-02$$$78.5 LiveBench1M context128K outputInkling SmallThinkingmachinesInkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency.2026-07-30$$1M context1M outputQwen 3.7 FlashAlibabaThe Qwen3.7 native vision-language Flash model series delivers a comprehensive upgrade over 3.6-Flash in multimodal understanding and agent execution. This model particularly excels in enhanced multimodal foundations with stronger universal object recognition, further improved real-world perception and spatial intelligence, significantly upgraded multimodal agent capabilities for Search Agent and CI Agent scenarios with more stable end-to-end task execution, as well as optimized multimodal coding for a smoother vibe coding experience.2026-07-28$991K context64K outputClaude Opus 5AnthropicClaude Opus 5 is the latest model in Anthropic's Opus family and a step-change improvement over Opus 4.8. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, and it is stronger per token across effort levels.2026-07-24$$$+80.1 LiveBench1M context128K outputGemini 3.5 Flash LiteGoogleGemini 3.5 Flash Lite features upgraded agentic capabilities, making the model ideal for subagents in complex workflows.2026-07-21$$63.9 LiveBench1M context65K outputGemini 3.6 FlashGoogleGemini 3.6 Flash delivers higher quality across coding, agentic workflows, and web development with reduced token consumption and fewer model calls compared to previous model iterations.2026-07-21$$73.6 LiveBench1M context64K outputKimi K3Moonshot AIKimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.2026-07-16$$$79.2 LiveBench1M context131.1K outputInklingThinkingmachinesInkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.2026-07-15$$71.9 LiveBench256K context256K outputMuse Spark 1.1MetaMuse Spark 1.1 is strongest at agentic performance, tool use, and computer use. It does well on long-running tasks with 1M token context window, can delegate execution to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser.2026-07-09$$$75.3 LiveBench1M context1M outputGPT 5.6 LunaOpenAIGPT-5.6 Luna is a fast, affordable GPT-5.6 model that brings strong capability at the lowest cost in the series.2026-07-09$73.6 LiveBench1.1M context128K outputGPT 5.6 SolOpenAIGPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 series, its most capable model for long-horizon agentic work across coding, biology, and cybersecurity.2026-07-09$$$81.1 LiveBench1.1M context128K outputGPT 5.6 TerraOpenAIGPT-5.6 Terra is a balanced GPT-5.6 model for everyday work, with performance comparable to the previous generation at half the cost.2026-07-09$$$77.9 LiveBench1.1M context128K outputHy3TencentTencent Hy3 is an open-source Mixture-of-Experts (MoE) large language model developed by Tencent's Hunyuan team. It has 295B total parameters, 21B active parameters, a 3.8B-parameter MTP layer, and a 256K context window. Hy3 is designed for agentic workflows, coding, reasoning, long-context understanding, tool-use scenarios, document analysis, and structured generation.
The model improves on Hy3 Preview with stronger tool-call reliability, better output-format stability, improved long-context retention, and lower hallucination rates in Tencent's internal evaluations.2026-07-06$256K context128K outputClaude Fable 5AnthropicClaude Fable 5 is a Mythos-class model with robust safeguards. It can handle long-running, complex, and asynchronous tasks where previous models would have needed more frequent check-ins.2026-07-01$$$+83.0 LiveBench1M context128K outputClaude Sonnet 5AnthropicSonnet 5 is an upgrade to Sonnet 4.6, with gains across agentic coding and professional work. It builds on the strengths of previous Sonnet models, bringing top-tier intelligence at Sonnet pricing for coding, agents, and everyday professional work at scale.2026-06-29$$$76.0 LiveBench1M context128K outputGLM 5.2Z.aiGLM-5.2 delivers powerful coding capabilities, usable 1M-context support, and continued strengths in long-horizon tasks.2026-06-16$$73.2 LiveBench1M context128K outputKimi K2.7 CodeMoonshot AIKimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.2026-06-12$$68.4 LiveBench256K context32.8K outputNemotron 3 UltraNVIDIAA 550B parameter (55B active) open reasoning model from NVIDIA, built for long-running agent workflows. It uses a hybrid Mamba-Transformer MoE architecture and supports a 1M token context window.2026-06-04$$51.8 LiveBench1M context65K outputQwen 3.7 PlusAlibabaAmong the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows.2026-06-02$$1M context64K outputClaude Opus 4.8AnthropicOpus 4.8 is a focused upgrade to Opus 4.7 and is Anthropic's best generally available model for coding, agentic tasks, and enterprise workflows. It builds on the strengths of previous Opus models with stronger performance on complex, multi-step coding tasks. Anthropic recommends using it on long-horizon coding and agentic tasks. It is also stronger on professional work, including document drafting, data analysis, and presentations.2026-05-28$$$+76.2 LiveBench1M context128K outputGemini 3.5 FlashGoogleGoogle's latest model, highly optimized for coding proficiency and parallel agentic execution loops. Defaults to medium thinking effort for faster and more cost-efficient responses.2026-05-19$$$74.6 LiveBench1M context64K outputGemini 3.1 Flash LiteGoogleGemini 3.1 Flash Lite outperforms 2.5 Flash Lite on overall quality and lands close to 2.5 Flash performance across key capability areas. It is a workhorse model for high-volume use cases, with improvements across audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion.2026-05-07$$61.7 LiveBench1M context65K outputMistral Medium LatestMistralMistral's frontier-class multimodal model optimized for agentic and coding use cases.2026-04-29$$$256K context256K outputGPT 5.5OpenAIGPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.2026-04-24$$$+80.2 LiveBench1M context128K outputDeepSeek V4 FlashDeepSeekDeepseek V4 Flash is an AI model available through Shortly AI.2026-04-23$65.5 LiveBench1M context384K outputDeepSeek V4 Flash 0731DeepSeekDeepseek V4 Flash 0731 is an AI model available through Shortly AI.2026-04-23$74.2 LiveBench1M context384K outputDeepSeek V4 ProDeepSeekDeepseek V4 Pro is an AI model available through Shortly AI.2026-04-23$$$71.6 LiveBench1M context1M outputMiMo M2.5XiaomiA native full-modal model supporting text, image, video, and audio understanding, with powerful Agent capabilities.2026-04-22$1.1M context131.1K outputMiMo V2.5 ProXiaomiMiMo V2.5 Pro delivers significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks. MiMo-V2.5-Pro is a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with a 1M-token context window.2026-04-22$$1.1M context131K outputKimi K2.6Moonshot AIKimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.2026-04-20$$70.5 LiveBench262K context262K outputClaude Opus 4.7AnthropicOpus 4.7 builds on the coding and agentic strengths of Opus 4.6 with stronger performance on complex, multi-step tasks and more reliable agentic execution. It also brings improved performance on knowledge work, from drafting documents to building presentations and analyzing data.2026-04-16$$$+76.5 LiveBench1M context128K outputGLM 5.1Z.aiGLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours—autonomously planning, executing, and improving itself throughout the process—ultimately delivering complete, engineering-grade results.2026-04-07$$$70.2 LiveBench202.8K context64K outputQwen 3.6 PlusAlibabaThe Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.2026-04-02$$68.9 LiveBench1M context64K outputGPT 5.4 MiniOpenAIGPT-5.4 Mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.2026-03-17$$66.4 LiveBench400K context128K outputGPT 5.4 NanoOpenAIGPT-5.4 Nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.2026-03-17$69.6 LiveBench400K context128K outputNVIDIA Nemotron 3 Super 120B A12BNVIDIANVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. It delivers up to 7x higher throughput, providing fast, cost-efficient inference for agentic tasks. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere.2026-03-11$32.5 LiveBench256K context32K outputGPT 5.4OpenAIGPT-5.4 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.2026-03-05$$$78.0 LiveBench1.1M context128K outputGemini 3.1 Pro PreviewGoogleThis model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows. Improvements highlighted include use cases for coding, multi-step function calling, planning, reasoning, deep knowledge tasks, and instruction following.2026-02-19$$$77.0 LiveBench1M context64K outputClaude Sonnet 4.6AnthropicClaude Sonnet 4.6 is the most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.2026-02-17$$$73.0 LiveBench1M context128K outputMiniMax M2.5MinimaxMiniMax-M2.5 is a SOTA large language model designed for real-world productivity. It is capable of handling the entire development process of various complex systems. It covers full-stack projects across multiple platforms including Web, Android, iOS, Windows, and Mac, encompassing server-side APIs, functional logic, and databases.2026-02-12$$60.1 LiveBench204.8K context131K outputGLM 5Z.aiGLM 5 is a frontier-class, general-purpose large language model optimized for complex systems engineering and long-horizon agentic tasks. It builds on the GLM 4.5 agent-centric lineage and is designed to support multi-step reasoning, math (including AIME-style benchmarks), advanced coding, and tool-augmented workflows, with long context support suitable for sophisticated agents and enterprise applications. Typical uses include autonomous agents for software engineering, data and systems troubleshooting, operations copilots, and high-end chat assistants that must break down complex tasks, call tools reliably, and reason over long sequences of instructions or documents.2026-02-12$$68.9 LiveBench202.8K context131.1K outputClaude Opus 4.6AnthropicOpus 4.6 is the world’s best model for coding and professional work, built to power agents that take on whole categories of real-world work. It excels across the entire SDLC, breaking through on hard problems, identifying complex bugs, and demonstrating deeper codebase understanding. It also delivers a step-change in knowledge work, with near-production-ready documents, presentations, and spreadsheets on the first pass.2026-02-05$$$+74.5 LiveBench1M context128K outputGPT 5.3 CodexOpenAIGPT-5.3-Codex advances both the frontier coding performance of GPT‑5.2-Codex and the reasoning and professional knowledge capabilities of GPT‑5.2, together in one model, which is also 25% faster. This enables it to take on long-running tasks that involve research, tool use, and complex execution.2026-02-05$$$71.6 LiveBench400K context128K outputKimi K2.5Moonshot AIkimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.2026-01-26$$69.1 LiveBench262.1K context262.1K outputGLM 4.7 FlashZ.aiGLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option. Beyond coding, it is also recommended for creative writing, translation, long-context tasks, and roleplay.2026-01-19$200K context131K outputMiniMax M2.1MinimaxMiniMax 2.1 is MiniMax's latest model, optimized specifically for robustness in coding, tool use, instruction following, and long-horizon planning.2025-12-23$$204.8K context131.1K outputGLM 4.7Z.aiGLM-4.7 is Z.ai’s latest flagship model, with major upgrades focused on two key areas: stronger coding capabilities and more stable multi-step reasoning and execution.2025-12-22$$62.7 LiveBench200K context120K outputGPT 5.2 CodexOpenAIGPT‑5.2-Codex is a version of GPT‑5.2 further optimized for agentic coding in Codex, including improvements on long-horizon work through context compaction, stronger performance on large code changes like refactors and migrations, improved performance in Windows environments, and significantly stronger cybersecurity capabilities.2025-12-18$$$74.0 LiveBench400K context128K outputGemini 3 FlashGoogleGoogle's most intelligent model built for speed, combining frontier intelligence with superior search and grounding.2025-12-17$$72.4 LiveBench1M context65K outputNemotron 3 Nano 30B A3BNVIDIANVIDIA Nemotron 3 Nano is an open reasoning model optimized for fast, cost-efficient inference. Built with a hybrid MoE and Mamba architecture and trained on NVIDIA-curated synthetic reasoning data, it delivers strong multi-step reasoning with stable latency and predictable performance for agentic and production workloads.2025-12-15$262.1K context262.1K outputGPT 5.2OpenAIGPT-5.2 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.2025-12-11$$$74.6 LiveBench400K context128K outputDevstral 2MistralAn enterprise-grade text model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents.2025-12-09$$256K context256K outputDevstral Small 2MistralOur open source model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents.2025-12-09$256K context256K outputNova 2 LiteAmazonNova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.2025-12-02$$1M context1M outputMinistral 14BMistralMinistral 3 14B is the largest model in the Ministral 3 family, offering state-of-the-art capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. Optimized for local deployment, it delivers high performance across diverse hardware, including local setups.2025-12-02$256K context256K outputMistral Large 3MistralMistral Large 3 2512 is Mistral’s most capable model to date. It has a sparse mixture-of-experts architecture with 41B active parameters (675B total).2025-12-02$$256K context256K outputDeepSeek V3.2DeepSeekDeepSeek-V3.2: Official successor to V3.2-Exp.2025-12-01$57.5 LiveBench128K context8K outputDeepSeek V3.2 ThinkingDeepSeekDeepSeek‑V3.2 from DeepSeek harmonizes high computational efficiency with superior reasoning and agent performance. It builds on three main techniques: DeepSeek Sparse Attention for long‑context efficiency, a scalable reinforcement learning framework, and a large‑scale agentic task synthesis pipeline. This model excels at long-context reasoning and agentic tasks, efficiently handling extended inputs while maintaining strong accuracy. Its sparse attention design enables it to process complex, multi-step workflows without excessive compute costs. Overall, DeepSeek‑V3.2 targets long‑context reasoning, tool‑using agents, and efficient deployment in production environments.2025-12-01$$66.2 LiveBench128K context8K outputClaude Opus 4.5AnthropicClaude Opus 4.5 is Anthropic’s latest model in the Opus series, meant for demanding reasoning tasks and complex problem solving. This model has improvements in general intelligence and vision compared to previous iterations. In addition, it is suited for difficult coding tasks and agentic workflows, especially those with computer use and tool use, and can effectively handle context usage and external memory files.2025-11-24$$$+72.6 LiveBench200K context64K outputGPT-5.1-CodexOpenAIGPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments.2025-11-12$$$71.4 LiveBench400K context128K outputGPT 5.1 Codex MiniOpenAIGPT-5.1 Codex mini is a smaller, faster, and cheaper version of GPT-5.1 Codex.2025-11-12$$65.4 LiveBench400K context128K outputGPT 5.1 ThinkingOpenAIAn upgraded version of GPT-5 that adapts thinking time more precisely to the question to spend more time on complex questions and respond more quickly to simpler tasks.2025-11-12$$$72.0 LiveBench400K context128K outputKimi K2 ThinkingMoonshot AIKimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.2025-11-06$$65.6 LiveBench216.1K context216.1K outputNvidia Nemotron Nano 12B V2 VLNVIDIAThe model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities.2025-10-28$131.1K context131.1K outputClaude Haiku 4.5AnthropicClaude Haiku 4.5 matches Sonnet 4's performance on coding, computer use, and agent tasks at substantially lower cost and faster speeds. It delivers near-frontier performance and Claude’s unique character at a price point that works for scaled sub-agent deployments, free tier products, and intelligence-sensitive applications with budget constraints.2025-10-15$$$50.9 LiveBench200K context64K outputGLM 4.6Z.aiAs the latest iteration in the GLM series, GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications.2025-09-30$$59.5 LiveBench200K context96K outputClaude Sonnet 4.5AnthropicClaude Sonnet 4.5 is the newest model in the Sonnet series, offering improvements and updates over Sonnet 4.2025-09-29$$$59.2 LiveBench1M context64K outputQwen3 VL 235B A22B ThinkingAlibabaQwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade.2025-09-23$$131.1K context32.8K outputQwen3 VL 235B A22B InstructAlibabaThe Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement.2025-09-23$$131.1K context129K outputGPT-5-CodexOpenAIGPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments.2025-09-15$$$78.2 LiveBench400K context128K outputQwen3 Next 80B A3B InstructAlibabaA new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507).2025-09-11$53.1 LiveBench131.1K context32.8K outputQwen3 Next 80B A3B ThinkingAlibabaA new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507).2025-09-11$54.9 LiveBench131.1K context32.8K outputDeepSeek V3.1DeepSeekDeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.2025-08-21$163.8K context128K outputNvidia Nemotron Nano 9B V2NVIDIANVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.\2025-08-18$131.1K context131.1K outputGPT-5OpenAIGPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.2025-08-07$$$78.8 LiveBench400K context128K outputGPT-5 miniOpenAIGPT-5 mini is a cost optimized model that excels at reasoning/chat tasks. It offers an optimal balance between speed, cost, and capability.2025-08-07$$66.2 LiveBench400K context128K outputGPT-5 nanoOpenAIGPT-5 nano is a high throughput model that excels at simple instruction or classification tasks.2025-08-07$53.6 LiveBench400K context128K outputGPT OSS 120BOpenAIExtremely capable general-purpose LLM with strong, controllable reasoning capabilities2025-08-05$50.4 LiveBench131.1K context131.1K outputGPT OSS 20BOpenAIA compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments.2025-08-05$131.1K context8.2K outputQwen 3 Coder 30B A3B InstructAlibabaEfficient coding specialist balancing performance with cost-effectiveness for daily development tasks while maintaining strong tool integration capabilities.2025-07-31$262.1K context8.2K outputQwen3 Coder 480B A35B InstructAlibabaQwen3-Coder-480B-A35B-Instruct is a cutting-edge open coding model from Qwen, matching Claude Sonnet’s performance in agentic programming, browser automation, and core development tasks.2025-07-22$$$61.7 LiveBench262.1K context65.5K outputQwen3 Coder NextAlibabaQwen3-Coder-Next is an open-weight language model built specifically for coding, with strong performance on large-scale software engineering and agentic coding benchmarks. It uses a hybrid Mixture-of-Experts architecture to offer high capability at relatively modest active parameter counts, improving efficiency for real-world deployments. The model is trained on diverse code and natural language data so it can handle tasks like code generation, refactoring, debugging, repository-level reasoning, and technical explanation across multiple programming languages. It is also optimized for tool use and function calling, making it suitable as the core of coding agents that interact with shells, editors, issue trackers, and other developer tools.2025-07-22$$256K context256K outputGemini 2.5 Flash LiteGoogleGemini 2.5 Flash-Lite is a balanced, low-latency model with configurable thinking budgets and tool connectivity (e.g., Google Search grounding and code execution). It supports multimodal input and offers a 1M-token context window.2025-06-17$47.8 LiveBench1M context65.5K outputClaude Sonnet 4AnthropicClaude Sonnet 4 balances impressive performance for coding with the right speed and cost for high-volume use cases: Coding: Handle everyday development tasks with enhanced performance-power code reviews, bug fixes, API integrations, and feature development with immediate feedback loops.2025-05-22$$$51.0 LiveBench1M context8.2K outputGPT-4.1 miniOpenAIGPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases.2025-05-14$$59.1 LiveBench1M context32.8K outputMistral Medium 3.1MistralMistral Medium 3 delivers frontier performance while being an order of magnitude less expensive. For instance, the model performs at or above 90% of Claude Sonnet 3.7 on benchmarks across the board at a significantly lower cost.2025-05-07$$128K context64K outputQwen3-14BAlibabaQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support2025-04-28$41K context16.4K outputQwen3-30B-A3BAlibabaQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support2025-04-28$39.0 LiveBench41K context16.4K outputQwen 3 32BAlibabaQwen3-32B is a world-class model with comparable quality to DeepSeek R1 while outperforming GPT-4.1 and Claude Sonnet 3.7. It excels in code-gen, tool-calling, and advanced reasoning, making it an exceptional model for a wide range of production use cases.2025-04-28$43.6 LiveBench128K context8.2K outputo3OpenAIOpenAI's o3 is their most powerful reasoning model, setting new state-of-the-art benchmarks in coding, math, science, and visual perception. It excels at complex queries requiring multi-faceted analysis, with particular strength in analyzing images, charts, and graphics.2025-04-16$$$80.7 LiveBench200K context100K outputo4-miniOpenAIOpenAI's o4-mini delivers fast, cost-efficient reasoning with exceptional performance for its size, particularly excelling in math (best-performing on AIME benchmarks), coding, and visual tasks.2025-04-16$$78.7 LiveBench200K context100K outputGPT-4.1OpenAIGPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.2025-04-14$$$63.0 LiveBench1M context32.8K outputGPT-4.1 nanoOpenAIGPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model.2025-04-14$46.6 LiveBench1M context32.8K outputLlama 4 Maverick 17B InstructMetaAs a general purpose LLM, Llama 4 Maverick contains 17 billion active parameters, 128 experts, and 400 billion total parameters, offering high quality at a lower price compared to Llama 3.3 70B.2025-04-05$54.4 LiveBench128K context8.2K outputLlama 4 Scout 17B InstructMetaLlama 4 Scout is the best multimodal model in the world in its class and is more powerful than our Llama 3 models, while fitting in a single H100 GPU. Additionally, Llama 4 Scout supports an industry-leading context window of up to 10M tokens.2025-04-05$128K context8.2K outputGemini 2.5 FlashGoogleGemini 2.5 Flash is a thinking model that offers great, well-rounded capabilities. It is designed to offer a balance between price and performance with multimodal support and a 1M token context window.2025-03-20$$53.2 LiveBench1M context65.5K outputGemini 2.5 ProGoogleGemini 2.5 Pro is our most advanced reasoning Gemini model, capable of solving complex problems. Gemini 2.5 Pro can comprehend vast datasets and challenging problems from different information sources, including text, audio, images, video, and even entire code repositories.2025-03-20$$$63.3 LiveBench1M context65.5K outputCommand ACohereCommand A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.2025-03-13$$$44.1 LiveBench256K context8K outputo3-miniOpenAIo3-mini is OpenAI's most recent small reasoning model, providing high intelligence at the same cost and latency targets of o1-mini.2025-01-31$$71.4 LiveBench200K context100K outputDeepSeek-R1DeepSeekDeepSeek-R1 provides customers a state-of-the-art reasoning model, optimized for general reasoning tasks, math, science, and code generation.2025-01-20$$$67.5 LiveBench128K context8.2K outputDeepSeek V3 0324DeepSeekDeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.2024-12-26$62.8 LiveBench163.8K context163.8K outputLlama 3.3 70B InstructMetaWhere performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation.2024-12-06$$128K context8.2K outputNova LiteAmazonA very low cost multimodal model that is lightning fast for processing image, video, and text inputs.2024-12-03$300K context8.2K outputNova MicroAmazonA text-only model that delivers the lowest latency responses at very low cost.2024-12-03$128K context8.2K outputNova ProAmazonA highly capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks.2024-12-03$$300K context8.2K outputMinistral 3BMistralA compact, efficient model for on-device tasks like smart assistants and local analytics, offering low-latency performance.2024-10-16$128K context4K outputMinistral 8BMistralA more powerful model with faster, memory-efficient inference, ideal for complex workflows and demanding edge applications.2024-10-16$128K context4K outputMistral SmallMistralMistral Small is the ideal choice for simple tasks that one can do in bulk - like Classification, Customer Support, or Text Generation. It offers excellent performance at an affordable price point.2024-09-17$32K context4K outputPixtral 12B 2409MistralA 12B model with image understanding capabilities in addition to text.2024-09-17$128K context4K outputLlama 3.1 70B InstructMetaAn update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.2024-07-23$$128K context8.2K outputLlama 3.1 8B InstructMetaAn update to Meta Llama 3 8B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.2024-07-23$128K context8.2K outputMistral Nemo 12BMistralA 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license.2024-07-18$128K context128K outputGPT-4o miniOpenAIGPT-4o mini from OpenAI is their most advanced and cost-efficient small model. It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast.2024-07-18$43.4 LiveBench128K context16.4K outputMistral CodestralMistralMistral's cutting-edge language model for coding released end of July 2025, Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation.2024-05-29$128K context4K outputGPT-4oOpenAIGPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API.2024-05-13$$$53.9 LiveBench128K context16.4K outputClaude 3 HaikuAnthropicClaude 3 Haiku is Anthropic's fastest, most compact model for near-instant responsiveness. It answers simple queries and requests with speed. Customers will be able to build seamless AI experiences that mimic human interactions. Claude 3 Haiku can process images and return text outputs, and features a 200K context window.2024-03-13$$35.4 LiveBench200K context4.1K outputGPT-3.5 TurboOpenAIOpenAI's most capable and cost effective model in the GPT-3.5 family optimized for chat purposes, but also works well for traditional completions tasks.2023-03-01$$33.2 LiveBench16.4K context4.1K output