Meta: Meta Llama 3 8B Instruct
novita
Lowest priceMeta-Llama-3-8B-Instruct is an instruction-tuned variant of Meta's 8-billion-parameter transformer, optimized for dialogue and task completion. Trained with RLHF on 15T+ tokens, it delivers strong performance in reasoning, coding, and multilingual tasks while maintaining efficiency for single-GPU deployment. Features improved instruction following and reduced hallucination compared to base models.
Tool Calling
Google: Gemma 3 12B IT
novita
Gemma 3 12B Instruct is a 12-billion-parameter multimodal model from Google, built from the same research as the Gemini models. It is instruction-tuned for dialogue and excels at text generation and image understanding tasks like question answering, summarization, and reasoning.
Tool CallingVisionLong Context
OpenAI: GPT Oss 20B
thinkingmachines
GPT-OSS-20B is OpenAI's efficiency-optimized 20-billion-parameter transformer featuring architectural innovations from GPT-4. Designed for accessible deployment, it delivers 85% of GPT-OSS-120B's capability at 6× lower resource requirements. Includes sparse attention mechanisms, constitutional AI safeguards, and deterministic output options.
Tool Calling
Qwen: Qwen3 Next 80B A3B Instruct
novita
Qwen3-Next-80B-A3B-Instruct is an 80B-parameter MoE model with 3B active parameters per token, optimized for instruction following and complex reasoning tasks. Features enhanced multilingual capabilities, advanced tool use, and extended context handling with improved efficiency through expert routing.
Tool CallingLong Context
Mistral AI: Mistral Nemo Instruct 2407
novita
Mistral-Nemo-Instruct-2407 is a 12-billion-parameter instruction-tuned language model developed jointly by Mistral AI and NVIDIA. It features a 128k token context window, strong multilingual capabilities, and excels at tasks like text generation, code generation, and reasoning. It uses the efficient Tekken tokenizer and is released under the Apache 2.0 license.
Tool CallingLong Context
NVIDIA: Nemotron-3 Nano 30B A3B
novita
NVIDIA Nemotron-3 Nano 30B A3B is NVIDIA's open reasoning model featuring a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture. It consists of 23 Mamba-2 and MoE layers, plus 6 Attention layers. The model is built for building specialized AI agents, chatbots, and RAG systems, providing high compute efficiency. It natively supports configurable reasoning depth through a 'thinking budget' and acts as a general-purpose reasoning and chat model intended for English, coding languages, and several other supported languages.
Tool CallingLong Context
DeepSeek: DeepSeek R1 Distill Qwen 14B
novita
DeepSeek-R1-Distill-Qwen-14B is a knowledge-distilled model combining DeepSeek-R1's reasoning capabilities with Qwen's multilingual strengths. Features enhanced Chinese-English performance and balanced task generalization at 14B scale with FP16 precision.
Tool Calling
GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.
Tool CallingVisionLong Context
Qwen: Qwen3 235B A22B Instruct 2507
novita
Qwen3-235B-A22B-Instruct-2507 is Alibaba Cloud's frontier mixture-of-experts model featuring 235B total parameters with 22 specialized experts and 4 active per token. Designed for superhuman reasoning and multilingual mastery, it achieves state-of-the-art performance across technical, creative, and analytical domains with 256K context handling. Optimized for distributed inference across H100 GPU clusters.
Tool CallingVisionLong Context
ByteDance: Seed 1.6 Flash
bytedance
ByteDance Seed 1.6 Flash is an ultra-fast multimodal deep-thinking model optimized for speed and cost-efficiency without compromising capability. It supports text, image, and video inputs along with tool use and function calling across an expansive 256K token context window, making it ideal for high-speed automated workflows and media-heavy processing.
Tool CallingVisionLong Context
DeepSeek: DeepSeek V4 Flash
novita
DeepSeek-V4-Flash is a highly efficient Mixture-of-Experts (MoE) language model featuring 158B total parameters with only 13B activated during inference. It utilizes a Hybrid Attention Architecture (combining Compressed Sparse Attention and Heavily Compressed Attention) to natively and efficiently support a 1-million-token context window. The model offers switchable reasoning modes (Non-think, Think High, and Think Max) and excels at high-speed generation for reasoning, coding, tool-calling, and agentic tasks.
Tool CallingLong Context
GPT-5 Nano is a lightweight, low-latency language model designed for cost-efficient text generation, chat, and tool-calling workloads.
Tool CallingLong Context
OpenAI: GPT 4.1 Nano
openai
GPT-4.1 Nano is a highly efficient 100B-parameter variant optimized for edge deployment and mobile devices. Features 128 experts with aggressive pruning and INT8 quantization, delivering 60% of GPT-4.1's capability while consuming 75% less power.
Tool CallingVision
OpenAI: GPT Oss 120B
novita
GPT-OSS-120B is OpenAI's open-source 120-billion-parameter transformer model featuring architectural innovations from GPT-4. Optimized for large-scale deployment with custom FP8 quantization, it delivers state-of-the-art reasoning capabilities while maintaining 3× better throughput than comparable models. Includes constitutional AI safeguards and deterministic output options.
Tool CallingLong Context
Google: Gemma 4 26B A4b
novita
Gemma 4 26B A4B is an efficient, open-weights Mixture-of-Experts (MoE) multimodal model from Google DeepMind. With only 3.8B active parameters out of 25.2B total, it runs nearly as fast as a 4B model while delivering performance close to the 31B dense model. Features a 256K token context window, supports interleaved text and image inputs, and excels at reasoning, coding, and agentic workflows.
Tool CallingVisionLong Context
Google: Gemma 4 31B
novita
Gemma 4 31B is a state-of-the-art, open-weights multimodal dense model from Google DeepMind. It features a 256K token context window, excels at reasoning, coding, and agentic workflows, and supports interleaved text and image inputs.
Tool CallingVisionLong Context
Qwen: Qwen3.8 Flash
openrouter
Qwen3.8-Flash is Qwen's high-speed multimodal reasoning model designed for coding, agentic workflows, visual understanding, long-context processing, and high-concurrency applications. It natively supports a 1M-token context window and can process text, images, video, lengthy documents, and large codebases. The model combines strong reasoning and generation capabilities with efficient inference, and supports built-in tools for agentic and developer workflows.
Tool CallingVisionLong Context
DeepSeek: DeepSeek V3.2
novita
DeepSeek-V3.2 is a general-purpose Mixture-of-Experts (MoE) large language model from DeepSeek AI that harmonizes high computational efficiency with superior reasoning and agent performance. It features 685 billion total parameters, supports a 164Ktoken context window, and is natively calibrated for tool use and function calling in complex agentic workflows.
Tool CallingLong Context
OpenAI: GPT 4O Mini
openai
GPT-4o Mini is a distilled 400B-parameter version of GPT-4o optimized for cost-efficient omnimodal applications. Maintains core multimodal capabilities with 70% of GPT-4o's performance at 40% of the computational cost, featuring real-time audio/image processing optimizations.
Tool CallingVisionLong Context
Qwen: Qwen 2.5 72B Instruct
novita
Qwen-2.5-72B-Instruct is Alibaba's flagship 72B parameter model featuring enhanced reasoning, 128K context, and advanced tool-calling capabilities. Excels in multilingual tasks, technical domains, and complex instruction following with improved safety alignment.
Tool CallingLong Context
GLM-4.6V is a vision-language model designed for cloud and high-performance clusters. It introduces native multimodal function calling, interleaved image-text generation, and advanced document understanding. It supports a 128K context window and achieves SoTA visual understanding among models of similar scale.
Tool CallingVisionLong Context
Microsoft: Wizardlm 2 8X22b
novita
WizardLM-2 8x22B is a 141-billion-parameter Mixture of Experts (MoE) model developed by Microsoft AI. It is the flagship model of the WizardLM-2 family, designed for state-of-the-art performance in complex chat, reasoning, multilingual tasks, and agent-like capabilities. It is built upon the Mixtral-8x22B architecture and trained using a fully AI-powered synthetic training system .
Tool Calling
DeepSeek: DeepSeek V3.1
novita
DeepSeek-V3.1 is a 1.3T-parameter Mixture-of-Experts model with 236B active parameters per token. It features enhanced reasoning capabilities, extended 1M token context window, and improved multilingual support with advanced tool usage and code generation capabilities.
Tool CallingLong Context
OpenAI: GPT 5.6 Luna
openai
GPT-5.6 Luna is the fastest, most cost-efficient model in the GPT-5.6 family. It is designed for high-volume, latency-sensitive workloads such as classification, data extraction, request routing, and draft generation. It shares the same 1.05M context window as its siblings
Tool CallingVisionLong Context
DeepSeek: DeepSeek V3 0324
novita
DeepSeek-V3-0324 is a 671B-parameter model with 37B active parameters per token. Features strong reasoning capabilities, 128K context window, and enhanced multilingual support with advanced tool usage and code generation.
Tool CallingLong Context
OpenAI: GPT 5.4 Nano
openai
GPT-5.4 Nano is the smallest and fastest model in the GPT-5.4 lineup, designed for ultra-low latency and low-cost API usage at high throughput. It is optimized for short-turn tasks like classification, extraction, ranking, and lightweight sub-agent work.
Tool CallingVisionLong Context
Qwen: Qwen3 Coder 480B A35B Instruct FP8
novita
Qwen3-Coder-480B-A35B-Instruct is a massive 480B-parameter MoE model specialized in code generation and technical problem-solving. Featuring 35B active parameters per token through expert routing, it delivers state-of-the-art coding assistance with FP8 quantization for efficient deployment.
Tool CallingLong Context
MiniMax: MiniMax M2.7
minimax
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.
Tool CallingLong Context
MiniMax: MiniMax M2.5
minimax
MiniMax M2.5 is an open-source MoE model with 230B total parameters (10B active) and a 200K token context window, released February 2026 under a Modified MIT license. SOTA in coding, agentic tool use, search, and office work — scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp. Trained across 200K+ real-world RL environments in 10+ programming languages.Matches Claude Opus 4.6 speed while completing SWE-Bench 37% faster than M2.1.
Tool CallingLong Context
MiniMax: MiniMax M3
novita
MiniMax M3 is an open model with sparse block-level attention designed for coding, agentic workflows, and multimodal chat. It supports a 1M token context window and operates natively across text, image, and video.
Tool CallingVisionLong Context
Baidu: Ernie 4.5 VL 424B A47B Base Pt
novita
ERNIE 4.5 VL is Baidu's 424-billion parameter vision-language Mixture of Experts model with 47B active parameters per token. It features advanced multimodal understanding, reasoning across text and images, and state-of-the-art performance in complex visual-language tasks.
Tool CallingVision
Meta: Muse Glimmer 30B
together_ai
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for local autonomous agentic workflows on consumer hardware. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a built-in perception encoder. It supports interleaved text and image input, controllable reasoning strength, and is highly compatible with agentic orchestration frameworks.
Tool CallingVisionLong Context
OpenAI: GPT 4.1 Mini
openai
GPT-4.1 Mini is a distilled 400B-parameter version of GPT-4.1 optimized for cost-efficient deployment. Features the same 256-expert MoE architecture with selective expert activation, delivering 80% of GPT-4.1's capability at 30% of the inference cost.
Tool CallingVisionLong Context
OpenAI's most compact and cost-effective GPT-5 model, released August 7, 2025. Designed for high-volume, latency-sensitive applications where the full capabilities of larger models are not required, while maintaining strong performance on common tasks.
Tool CallingLong Context
ByteDance: Seed 1.8
bytedance
ByteDance Seed 1.8 is a large language model optimized for long-context reasoning and generation with a 256K context window. It excels in complex coding, multi-turn dialogue, and multi-step agentic tasks while robustly supporting vision inputs and parallel tool calling.
Tool CallingVisionLong Context
ByteDance: Seed 1.6
bytedance
A model from ByteDance featuring adaptive thinking (AdaCoT) to control reasoning length. It was developed from Seed1.5-Thinking with greater training compute and integrated VLM capabilities. Supports text and image inputs with a 256K context window
Tool CallingVisionLong Context
GLM-4.6 is a large language model with improvements over GLM-4.5, featuring a 200K context window, superior coding performance, advanced reasoning, and stronger agentic capabilities in tool use and search-based agents. It also offers refined writing and more natural role-playing.
Tool CallingLong Context
GLM-4.7 is a large language model focused on coding, agentic tasks, and complex reasoning. It introduces Interleaved Thinking, Preserved Thinking, and Turn-level Thinking for improved stability and performance in multi-turn, long-horizon tasks. It shows significant gains over its predecessor in multilingual agentic coding, terminal-based tasks, and tool use.
Tool CallingLong Context
Moonshot AI: Kimi K2 Instruct
novita
Kimi-K2-Instruct is Moonshot AI's flagship 100-billion-parameter instruction-tuned model featuring a hybrid dense-MoE architecture with 200K context handling. Optimized for complex reasoning, multilingual dialogue, and long-form document understanding. Specializes in Chinese and English technical domains with enhanced safety alignment.
Tool CallingLong Context
ByteDance: Dola Seed 2.1 Turbo
bytedance
Dola-Seed 2.1 Turbo is a next-generation multimodal large model by ByteDance designed for real-world productivity, high-value production tasks in enterprise R&D, and large-scale agentic scenarios. It features comprehensive upgrades in coding engineering delivery, long-horizon agent task execution, and multimodal understanding, along with stronger autonomous planning and dynamic self-repair capabilities.
Tool CallingVisionLong Context
Moonshot AI: Kimi K2 0905
novita
Kimi-K2-0905 is a 72B-parameter multimodal language model optimized for long-context understanding and complex reasoning. Features enhanced multilingual capabilities, advanced tool usage, and strong performance in mathematical and logical reasoning tasks with extended context handling.
Tool CallingVisionLong Context
Qwen: Qwen3 235B A22B Thinking 2507
novita
Qwen3-235B-A22B-Thinking-2507 is Alibaba Cloud's cognitive architecture model specializing in advanced reasoning, theory of mind simulation, and multi-agent collaboration. Built on the Qwen3-235B MoE foundation, it features specialized 'cognitive experts' and a neuro-symbolic execution engine for human-like problem-solving. Excels at strategic planning, counterfactual reasoning, and complex system modeling with 256K context retention.
Tool CallingLong Context
Qwen3.8-27B is a compact, deployment-friendly dense multimodal model from Alibaba's Qwen3.8 family. Natively supporting text, image, and video inputs, it handles tasks ranging from STEM diagrams and document understanding to long-form video analysis. It features flexible thinking control with adjustable reasoning effort and is designed for autonomous planning, coding, professional work, multimodal reasoning, and long-horizon agentic workflows.
Tool CallingVisionLong Context
ByteDance: Dola Seed 2.0 Code
bytedance
Dola-Seed 2.0 Code is ByteDance's dedicated coding model optimized for enterprise-level programming needs. It excels in agentic task planning with rational decomposition and dynamic self-repair, while inheriting the Seed 2.0 family's robust multimodal perception to read charts, screenshots, and diagrams for contextually accurate code generation.
Tool CallingVisionLong Context
Moonshot AI: Kimi K2.5
novita
Kimi K2.5 is a large open-weight MoE model with 1.1 trillion parameters, known for strong vision capabilities via MoonViT-3D and agent swarm orchestration.
Tool CallingVisionLong Context
Grok 4.20 Reasoning is a high-performance model by xAI designed for complex logic and math, scientific and technical analysis, multi-step investigations, and high-stakes tasks where accuracy matters most. It incorporates a 'thinking' mechanism before responding and offers strict prompt adherence alongside agentic tool calling.
Tool CallingVisionLong Context
Grok 4.3 is a flagship reasoning model from xAI leveraging a four-agent deliberation system where internal sub-agents debate the best response. It features native video input and leverages real-time data streams from the X platform for current-events knowledge. It is optimized for long-document analysis, conversational research, and complex multi-step agentic workflows.
Tool CallingVisionLong Context
Moonshot AI: Kimi K2.6
novita
Kimi-K2.6 is an open-source native multimodal agentic model developed by Moonshot AI, supporting long-horizon coding capabilities and agentic task orchestration scaling to 300 sub-agents.
Tool CallingVisionLong Context
GLM-5 is a 754B-parameter Mixture-of-Experts model from Z.ai. It features DeepSeek Sparse Attention (DSA) for efficient long-context handling, leading open-source models in long-horizon planning and agentic tasks.
Tool CallingLong Context
NVIDIA: Nemotron-3 Ultra 550B A55B
together_ai
NVIDIA Nemotron-3 Ultra 550B A55B is a frontier-scale large language model from NVIDIA, designed for strong agentic, reasoning, and conversational capabilities. It employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2, MoE, and Attention layers, alongside Multi-Token Prediction (MTP) for faster inference. With a 1-million-token context window and a configurable reasoning mode, it excels at complex autonomous agents, long-context analysis, and deep research workflows.
Tool CallingLong Context
DeepSeek: DeepSeek V4 Pro
novita
DeepSeek-V4-Pro is a flagship Mixture-of-Experts (MoE) language model with 862B total parameters (49B activated) and a 1-million-token context window. It features a Hybrid Attention Architecture combining Compressed Sparse Attention and Heavily Compressed Attention for efficient long-context processing. With multiple distinct reasoning modes (such as Think Max for maximum reasoning effort), it is built for advanced mathematical reasoning, software engineering, tool use scenarios, and complex long-horizon agentic workflows.
Tool CallingLong Context
Moonshot AI: Kimi K2.7 Code
novita
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
Tool CallingVisionLong Context
OpenAI: GPT 5.4 Mini
openai
GPT-5.4 Mini is a compact, cost-efficient model designed for reliable performance across high-volume, everyday AI workloads. It offers fast, efficient reasoning, strong multimodal capabilities, and approaches the performance of the larger GPT-5.4 model on coding and agentic tasks while running more than 2x faster than its predecessor.
Tool CallingVisionLong Context
Muse Spark 1.3 is Meta's natively multimodal reasoning model optimized for long-horizon agentic workflows, advanced coding, tool use, computer interaction, and complex multi-step tasks. It is designed to sustain multiple workflows within long conversations, generate and manage context across conflicting sources, identify and correct gaps in its plans, and efficiently complete real-world engineering and knowledge-work tasks. Compared with Muse Spark 1.2, it improves coding and agentic performance while using fewer tool calls and fewer tokens.
Tool CallingVisionLong Context
GLM-5.3 is Z.ai's advanced Mixture-of-Experts model, building upon the GLM-5.2 base architecture with significantly scaled post-training. It is designed specifically for complex coding, long-horizon tasks, and autonomous agent workflows. The model sets a new state-of-the-art for open models in agentic coding and demonstrates strong emergent cybersecurity capabilities, excelling in vulnerability discovery and exploitation chaining.
Tool CallingLong Context
string
Tool CallingLong Context
GLM-5.2 is Z.ai's latest flagship open-weights MoE model engineered specifically to dominate long-horizon autonomous coding and engineering tasks. It features a solid 1-million-token context window, multiple thinking effort levels, and an improved IndexShare architecture that significantly boosts reasoning and agentic performance.
Tool CallingLong Context
Anthropic: Claude Sonnet 5
anthropic
Claude Sonnet 5 is the most agentic Sonnet model yet, released by Anthropic on June 30, 2026. It offers a 1M token context window and 128k max output tokens. It is designed for complex, multi-step agentic tasks like coding, tool use, and autonomous planning, achieving performance close to the more expensive Opus 4.8 model. It features adaptive thinking enabled by default and a new tokenizer that produces approximately 30% more tokens for the same text compared to Sonnet 4.6.
Tool CallingVisionLong Context
Qwen3.8-Max is Alibaba's 2.4-trillion-parameter flagship multimodal large language model. Designed for long-horizon professional workflows, it excels in full-stack coding, agentic development, data analysis, and office automation. It processes text, images, video, and documents natively, and is positioned as a leading frontier model.
Tool CallingVisionLong Context
Grok 4.5 is xAI's frontier model built for coding, agentic tasks, and knowledge work. Trained on datasets spanning science, engineering, math, and code (in partnership with Cursor), it features a 1.5 trillion parameter architecture and exceeds comparable models in real-world software engineering tasks.
Tool CallingVisionLong Context
Grok 4.6 is an incremental flagship upgrade over Grok 4.5, utilizing the same 1.5 trillion parameter foundation but enhanced with superior supervised fine-tuning and reinforcement learning via the Grok Build harness. It brings frontier intelligence to long-running AI agents, complex multi-step tasks, and software development, offering aggressive verification and highly competitive cost-per-task efficiency.
Tool CallingVisionLong Context
Thinking Machines: Inkling Small
thinkingmachines
Inkling-Small is a general-purpose multimodal Mixture-of-Experts (MoE) model featuring 266B total parameters with only 12B active during inference. Natively processing text, image, and audio inputs, it supports a 256K token context window and variable thinking effort. Despite being a quarter of the size of the flagship Inkling model, it matches or exceeds its predecessor on key coding and agentic tasks while requiring significantly less compute.
Tool CallingVisionLong Context
GPT-4.1 is OpenAI's 1.6 trillion parameter multimodal reasoning system featuring a 256-expert MoE architecture with enhanced tool integration and world knowledge grounding. Optimized for enterprise applications with improved factual accuracy and reduced hallucination.
Tool CallingVisionLong Context
A large, multimodal language model with advanced reasoning capabilities and support for tools. It supports the new OpenAI Responses API with features like reasoning effort control.
Tool CallingVisionLong Context
OpenAI's GPT-5.1 model, released November 13, 2025. Features a strong multimodal and reasoning capabilities, and is optimized for complex agentic workflows. It introduced significant improvements in coding, reasoning, and tool use over previous generations.
Tool CallingVisionLong Context
GPT-4o (Omnimodal) is OpenAI's 1.2 trillion parameter multimodal foundation model featuring unified input processing across text, vision, and audio. Optimized for real-time interaction with enhanced reasoning and cross-modal understanding capabilities.
Tool CallingVisionLong Context
OpenAI: GPT 5.6 Terra
openai
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, the balanced workhorse tier of the GPT-5.6 family. Delivers highly reliable performance for everyday professional workflows, code generation, and general-purpose agentic tasks at a lower cost than Sol.
Tool CallingVisionLong Context
The standard 'Thinking' variant of OpenAI's GPT-5.2 model, released Dec 11, 2025. Excels at complex reasoning, coding, and multi-step projects. Optimized for professional knowledge work and serves as the balanced option between speed and capability in the GPT-5.2 family.
Tool CallingVisionLong Context
GPT-5.4 is OpenAI's flagship frontier model built for sustained, multi-step reasoning with reliable follow-through. It excels at agentic workflows, research, document analysis, and powering complex internal tools with robust instruction following.
Tool CallingVisionLong Context
Anthropic: Claude Sonnet 4 6
anthropic
Claude Sonnet 4.6 is Anthropic's balanced production model, offering near-Opus level intelligence at a much lower cost. It features a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work, and design, along with an adaptive thinking capability.
Tool CallingVisionLong Context
Moonshot AI: Kimi K3
novita
Kimi’s most capable flagship model to date, with 2.8 trillion parameters. Built for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. The world’s first open-source model in the 3-trillion-parameter class.
Tool CallingVisionLong Context
OpenAI: GPT 5.6 Sol
openai
GPT-5.6 Sol is the flagship model in the GPT-5.6 family, designed for the most complex reasoning, agentic coding, and deep research tasks. It features a 1.05M token context window and excels in long-horizon, multi-step workflows. It supports advanced features like programmatic tool calling, multi-agent orchestration, and a 'max' reasoning effort
Tool CallingVisionLong Context
Thinking Machines: Inkling
thinkingmachines
Inkling is a flagship open-weights multimodal Mixture-of-Experts (MoE) model from Thinking Machines Lab featuring 975B total parameters with 41B active per token. It natively processes text, image, and audio inputs and supports a 256K context window. The model offers a controllable 'thinking-effort' dial to balance compute cost against reasoning quality, excelling across coding, agentic workflows, and general reasoning tasks.
Tool CallingVisionLong Context
Anthropic: Claude Opus 4 6
anthropic
Claude Opus 4.6 brings Anthropic's advanced reasoning capabilities to high-stakes workflows. With full adaptive thinking support and a 1-million-token context window, it excels at complex tasks across coding, cybersecurity, financial analysis, and large-scale enterprise agents.
Tool CallingVisionLong Context
Anthropic: Claude Opus 4 7
anthropic
Claude Opus 4.7 excels in autonomous multi-step engineering and long-horizon agentic workflows. It features an Effort Control layer for self-correction and advanced visual intelligence, making it an ideal choice for software engineering, legal, and financial analyses.
Tool CallingVisionLong Context
Anthropic: Claude Opus 4 8
anthropic
Claude Opus 4.8 is a workhorse flagship model designed for high-end reasoning and long-horizon agentic coding. It offers exceptional consistency and autonomy to sustain work on complex, long-running tasks.
Tool CallingVisionLong Context
Anthropic: Claude Opus 5
anthropic
Claude Opus 5 is Anthropic's flagship model designed for complex agentic coding and enterprise work, delivering intelligence close to Claude Fable 5 at half the price. It features a 1M token context window, adaptive thinking, and state-of-the-art performance on coding and knowledge work evaluations. It is optimized for bounded, complex tasks
Tool CallingVisionLong Context
GPT-5.5 is a highly capable and efficient model featuring dynamic routing between 'Instant' and 'Thinking' modes. It provides improvements in conceptual clarity, scientific research ability, accuracy during knowledge work, and polished conversational tones.
Tool CallingVisionLong Context
Anthropic: Claude Fable 5
anthropic
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins.
Tool CallingVisionLong Context
Anthropic: Claude Fable 5.1
anthropic
Claude Fable 5.1 is Anthropic's frontier intelligence model designed for demanding reasoning, long-horizon agentic work, advanced software engineering, multistep research, and complex enterprise workflows. It is built for tasks that can span hours and multiple applications, with strong planning, tool orchestration, failure recovery, coding, document understanding, and vision capabilities. The model supports a 1M-token context window, adaptive reasoning that is always enabled, and up to 128K output tokens.
Tool CallingVisionLong Context
OpenAI: GPT 6 Astra
openai
GPT-6 Astra is OpenAI's flagship model, released 3 September 2026. It costs $10 per million input tokens and $50 per million output, with a 1,050,000-token context window and 128,000 max output tokens. It accepts text and images, and supports streaming, function calling and structured outputs.
Tool CallingVisionLong Context
OpenAI's flagship GPT-5 Pro model, released October 6, 2025. A highly capable multimodal model featuring advanced reasoning, native tool use, and strong performance on complex professional and coding tasks. It introduced a structured 'reasoning effort' parameter for controlling response depth.
Tool CallingVisionLong Context
OpenAI: GPT 5.2 Pro
openai
OpenAI's flagship GPT-5.2 Pro model, Designed for professional knowledge work and agentic workflows, exclusive 'xhigh' reasoning effort, and top-tier performance on complex reasoning, coding, and scientific benchmarks.
Tool CallingVisionLong Context
OpenAI: GPT 5.4 Pro
openai
GPT-5.4 Pro offers deeper, higher-reliability reasoning for complex production scenarios. It is engineered for high-stakes agentic workflows, long-form analysis and synthesis, complex planning, and advanced reasoning tasks where accuracy is paramount.
Tool CallingVisionLong Context
OpenAI: GPT 5.5 Pro
openai
GPT-5.5 Pro is the highest-capability model in the 5.5 lineup, optimized for the hardest tasks and long-running workflows. It uses significant compute to 'think' harder before answering, making it ideal for deep research, highly complex coding, and extensive multi-step problem solving.
Tool CallingVisionLong Context
Anthropic: Claude 3 Haiku 20240307
6ce62836a581464f8ea1e6ee1d397a99
Claude 3 Haiku is Anthropic's fastest and most compact model in the Claude 3 family, optimized for speed and cost-effectiveness while maintaining strong performance for everyday tasks and high-volume operations.
Tool CallingVisionLong Context
Anthropic: Claude Haiku 4 5 20251001
1be62feac9264117aee04eb7cdc868f2
Claude Haiku 4.5 is Anthropic's fastest and most efficient model in the Claude 4 family, optimized for speed and cost-effectiveness while maintaining high quality performance for everyday tasks.
Tool CallingVision
Anthropic: Claude Opus 4 1 20250805
f0f29ca2cd6c408c817f31623e97e73f
Claude Opus 4.1 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning and analysis tasks. Part of the earlier Claude 4 generation alongside Claude Opus 4.
Tool CallingVisionLong Context
Anthropic: Claude Opus 4 20250514
9d44f0f5167143d6a783d42ca0ba64e4
Claude Opus 4 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning, analysis, and challenging tasks requiring deep understanding.
Tool CallingVisionLong Context
Anthropic: Claude Opus 4 5 20251101
14991ae9ac7d45298b4e7e05db98b4fc
Claude Opus 4.5 is Anthropic's most capable model in the Claude 4 family, offering superior performance for complex tasks requiring deep reasoning and analysis.
Tool CallingVision
Anthropic: Claude Sonnet 4 20250514
91cad413eb7148658649629510425980
Claude Sonnet 4 is a balanced model in the Claude 4 family, offering a strong combination of intelligence, speed, and efficiency for a wide range of tasks including reasoning, coding, and analysis.
Tool CallingVisionLong Context
Anthropic: Claude Sonnet 4 5 20250929
a40f171ceba14d75b2bf1749904d0965
Claude Sonnet 4.5 is Anthropic's smartest model and the flagship of the Claude 4 family, offering the best balance of intelligence, speed, and efficiency for everyday use. It excels at complex reasoning, coding, and analysis tasks.
Tool CallingVision
DeepSeek: DeepSeek R1 Distill Llama 8B
24c0b83803ff4ac8b32693aee54098e8
DeepSeek-R1-Distill-Llama-8B combines DeepSeek-R1 knowledge distillation with Llama architecture. Available in FP8 (H100+ only) and FP16 quantization, delivering efficient performance with 32K context.
Tool Calling
Sarvam AI: Sarvam 105B
sarvam
Sarvam-105B is an advanced Mixture-of-Experts (MoE) model with 10.3B active parameters, designed for superior performance across a wide range of complex tasks. It is highly optimized for complex reasoning, with particular strength in agentic tasks, mathematics, and coding.
Tool CallingLong Context