meta-llama/Meta-Llama-3-8B-Instruct
novita
Lowest priceMeta-Llama-3-8B-Instruct is an instruction-tuned variant of Meta's 8-billion-parameter transformer, optimized for dialogue and task completion. Trained with RLHF on 15T+ tokens, it delivers strong performance in reasoning, coding, and multilingual tasks while maintaining efficiency for single-GPU deployment. Features improved instruction following and reduced hallucination compared to base models.
Tool Calling
google/gemma-3-12b-it
novita
Gemma 3 12B Instruct is a 12-billion-parameter multimodal model from Google, built from the same research as the Gemini models. It is instruction-tuned for dialogue and excels at text generation and image understanding tasks like question answering, summarization, and reasoning.
Tool CallingVisionLong Context
Qwen/Qwen3-Next-80B-A3B-Instruct
novita
Qwen3-Next-80B-A3B-Instruct is an 80B-parameter MoE model with 3B active parameters per token, optimized for instruction following and complex reasoning tasks. Features enhanced multilingual capabilities, advanced tool use, and extended context handling with improved efficiency through expert routing.
Tool CallingLong Context
mistralai/Mistral-Nemo-Instruct-2407
novita
Mistral-Nemo-Instruct-2407 is a 12-billion-parameter instruction-tuned language model developed jointly by Mistral AI and NVIDIA. It features a 128k token context window, strong multilingual capabilities, and excels at tasks like text generation, code generation, and reasoning. It uses the efficient Tekken tokenizer and is released under the Apache 2.0 license.
Tool CallingLong Context
GPT-OSS-20B is OpenAI's efficiency-optimized 20-billion-parameter transformer featuring architectural innovations from GPT-4. Designed for accessible deployment, it delivers 85% of GPT-OSS-120B's capability at 6× lower resource requirements. Includes sparse attention mechanisms, constitutional AI safeguards, and deterministic output options.
Tool Calling
deepseek/deepseek-r1-distill-qwen-14b
novita
DeepSeek-R1-Distill-Qwen-14B is a knowledge-distilled model combining DeepSeek-R1's reasoning capabilities with Qwen's multilingual strengths. Features enhanced Chinese-English performance and balanced task generalization at 14B scale with FP16 precision.
Tool Calling
Qwen/Qwen3-235B-A22B-Instruct-2507
novita
Qwen3-235B-A22B-Instruct-2507 is Alibaba Cloud's frontier mixture-of-experts model featuring 235B total parameters with 22 specialized experts and 4 active per token. Designed for superhuman reasoning and multilingual mastery, it achieves state-of-the-art performance across technical, creative, and analytical domains with 256K context handling. Optimized for distributed inference across H100 GPU clusters.
Tool CallingVisionLong Context
Gemma2 9B-IT is Google's lightweight 9B-parameter instruction-tuned model optimized for responsible deployment. Features enhanced reasoning, multilingual support, and hardware-efficient design for accessible AI applications with built-in safety filters and tool integration capabilities.
Tool Calling
Gemma2 9B-IT is Google's lightweight 9.24B-parameter instruction-tuned model optimized for responsible deployment. Features enhanced reasoning, multilingual support, and hardware-efficient design for accessible AI applications with built-in safety filters and tool integration capabilities.
Tool Calling
meta-llama/llama-4-scout-17b-16e-instruct
groq
Llama-4-Scout-17B-16E is Meta's efficiency-optimized MoE model featuring 16 experts with ~2B active parameters per token. Designed for exploratory tasks with enhanced web navigation, research synthesis, and information discovery capabilities at low resource cost.
Tool CallingLong Context
GPT-5 Nano is a lightweight, low-latency language model designed for cost-efficient text generation, chat, and tool-calling workloads.
Tool CallingLong Context
GPT-4.1 Nano is a highly efficient 100B-parameter variant optimized for edge deployment and mobile devices. Features 128 experts with aggressive pruning and INT8 quantization, delivering 60% of GPT-4.1's capability while consuming 75% less power.
Tool CallingVision
meta-llama/llama-3.3-70b-instruct
groq
Llama-3.3-70B-Instruct is Meta's flagship 70B parameter model featuring enhanced reasoning, 128K context, and enterprise-grade instruction following. Represents a significant evolution over Llama 3.1 with improved tool integration, safety alignment, and complex task handling.
Tool CallingLong Context
google/gemma-4-26B-A4B
novita
Gemma 4 26B A4B is an efficient, open-weights Mixture-of-Experts (MoE) multimodal model from Google DeepMind. With only 3.8B active parameters out of 25.2B total, it runs nearly as fast as a 4B model while delivering performance close to the 31B dense model. Features a 256K token context window, supports interleaved text and image inputs, and excels at reasoning, coding, and agentic workflows.
Tool CallingVisionLong Context
Gemma 4 31B is a state-of-the-art, open-weights multimodal dense model from Google DeepMind. It features a 256K token context window, excels at reasoning, coding, and agentic workflows, and supports interleaved text and image inputs.
Tool CallingVisionLong Context
qwen/qwen3-30b-a3b-fp8
novita
Qwen3-30B-A3B-FP8 is a 30B-parameter MoE model with 3B active parameters per token, optimized with FP8 quantization. Balances high performance with practical deployment requirements across reasoning, multilingual, and coding tasks.
Tool CallingLong Context
openai/gpt-oss-120b
novita
GPT-OSS-120B is OpenAI's open-source 120-billion-parameter transformer model featuring architectural innovations from GPT-4. Optimized for large-scale deployment with custom FP8 quantization, it delivers state-of-the-art reasoning capabilities while maintaining 3× better throughput than comparable models. Includes constitutional AI safeguards and deterministic output options.
Tool CallingLong Context
GPT-4o Mini is a distilled 400B-parameter version of GPT-4o optimized for cost-efficient omnimodal applications. Maintains core multimodal capabilities with 70% of GPT-4o's performance at 40% of the computational cost, featuring real-time audio/image processing optimizations.
Tool CallingVisionLong Context
qwen/qwen-2.5-72b-instruct
novita
Qwen-2.5-72B-Instruct is Alibaba's flagship 72B parameter model featuring enhanced reasoning, 128K context, and advanced tool-calling capabilities. Excels in multilingual tasks, technical domains, and complex instruction following with improved safety alignment.
Tool CallingLong Context
meta-llama/llama-4-maverick-17b-128e-instruct
groq
Llama 4 Maverick 17B-128E Instruct is Meta's 17B-parameter MoE instruction model featuring 128 experts with 4 active per token. Optimized for complex task decomposition and tool-integrated reasoning with constitutional AI safety, delivering expert-level performance at accessible computational requirements.
Tool CallingLong Context
DeepSeek-Chat is a 67B-parameter conversational AI optimized for helpful, honest, and engaging dialogue. Features instruction-following capabilities, emotional intelligence, and context-aware responses with safety alignment across 30+ conversation domains.
Tool CallingLong Context
microsoft/WizardLM-2-8x22B
novita
WizardLM-2 8x22B is a 141-billion-parameter Mixture of Experts (MoE) model developed by Microsoft AI. It is the flagship model of the WizardLM-2 family, designed for state-of-the-art performance in complex chat, reasoning, multilingual tasks, and agent-like capabilities. It is built upon the Mixtral-8x22B architecture and trained using a fully AI-powered synthetic training system .
Tool Calling
deepseek-ai/DeepSeek-V3.1
novita
DeepSeek-V3.1 is a 1.3T-parameter Mixture-of-Experts model with 236B active parameters per token. It features enhanced reasoning capabilities, extended 1M token context window, and improved multilingual support with advanced tool usage and code generation capabilities.
Tool CallingLong Context
deepseek-ai/DeepSeek-V3-0324
novita
DeepSeek-V3-0324 is a 671B-parameter model with 37B active parameters per token. Features strong reasoning capabilities, 128K context window, and enhanced multilingual support with advanced tool usage and code generation.
Tool CallingLong Context
GPT-5.4 Nano is the smallest and fastest model in the GPT-5.4 lineup, designed for ultra-low latency and low-cost API usage at high throughput. It is optimized for short-turn tasks like classification, extraction, ranking, and lightweight sub-agent work.
Tool CallingVisionLong Context
Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
novita
Qwen3-Coder-480B-A35B-Instruct is a massive 480B-parameter MoE model specialized in code generation and technical problem-solving. Featuring 35B active parameters per token through expert routing, it delivers state-of-the-art coding assistance with FP8 quantization for efficient deployment.
Tool CallingLong Context
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.
Tool CallingLong Context
MiniMax M2.5 is an open-source MoE model with 230B total parameters (10B active) and a 200K token context window, released February 2026 under a Modified MIT license. SOTA in coding, agentic tool use, search, and office work — scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp. Trained across 200K+ real-world RL environments in 10+ programming languages.Matches Claude Opus 4.6 speed while completing SWE-Bench 37% faster than M2.1.
Tool CallingLong Context
MiniMaxAI/MiniMax-M3
novita
MiniMax M3 is an open model with sparse block-level attention designed for coding, agentic workflows, and multimodal chat. It supports a 1M token context window and operates natively across text, image, and video.
Tool CallingVisionLong Context
baidu/ERNIE-4.5-VL-424B-A47B-Base-PT
novita
ERNIE 4.5 VL is Baidu's 424-billion parameter vision-language Mixture of Experts model with 47B active parameters per token. It features advanced multimodal understanding, reasoning across text and images, and state-of-the-art performance in complex visual-language tasks.
Tool CallingVision
deepseek-r1-distill-llama-70b
groq
DeepSeek-R1-Distill-Llama-70B is a hybrid 70B-parameter model combining DeepSeek's reasoning capabilities with Llama-70B's architectural efficiency. Knowledge-distilled from DeepSeek-R1 to Llama-3 architecture, delivering 92% of R1's performance with 60% lower inference cost and enhanced tool integration.
Tool CallingLong Context
GPT-4.1 Mini is a distilled 400B-parameter version of GPT-4.1 optimized for cost-efficient deployment. Features the same 256-expert MoE architecture with selective expert activation, delivering 80% of GPT-4.1's capability at 30% of the inference cost.
Tool CallingVisionLong Context
OpenAI's most compact and cost-effective GPT-5 model, released August 7, 2025. Designed for high-volume, latency-sensitive applications where the full capabilities of larger models are not required, while maintaining strong performance on common tasks.
Tool CallingLong Context
deepseek-reasoner
deepseek
DeepSeek-Reasoner is a 340B-parameter hybrid MoE model specialized in multi-step logical reasoning, causal inference, and problem decomposition. Features chain-of-thought verification, uncertainty quantification, and tool-integrated reasoning across mathematical, scientific, and real-world decision-making domains.
Tool CallingLong Context
moonshotai/Kimi-K2-Instruct
novita
Kimi-K2-Instruct is Moonshot AI's flagship 100-billion-parameter instruction-tuned model featuring a hybrid dense-MoE architecture with 200K context handling. Optimized for complex reasoning, multilingual dialogue, and long-form document understanding. Specializes in Chinese and English technical domains with enhanced safety alignment.
Tool CallingLong Context
moonshotai/kimi-k2-0905
novita
Kimi-K2-0905 is a 72B-parameter multimodal language model optimized for long-context understanding and complex reasoning. Features enhanced multilingual capabilities, advanced tool usage, and strong performance in mathematical and logical reasoning tasks with extended context handling.
Tool CallingVisionLong Context
deepseek-ai/DeepSeek-Prover-V2-671B
novita
DeepSeek-Prover-v2-671b is a 671B-parameter hybrid dense-MoE model specialized in mathematical theorem proving and formal reasoning, featuring 128 experts with 4-6 active per token. Optimized with FP8 quantization for efficient large-scale inference.
Qwen/Qwen3-235B-A22B-Thinking-2507
novita
Qwen3-235B-A22B-Thinking-2507 is Alibaba Cloud's cognitive architecture model specializing in advanced reasoning, theory of mind simulation, and multi-agent collaboration. Built on the Qwen3-235B MoE foundation, it features specialized 'cognitive experts' and a neuro-symbolic execution engine for human-like problem-solving. Excels at strategic planning, counterfactual reasoning, and complex system modeling with 256K context retention.
Tool CallingLong Context
moonshotai/Kimi-K2.5
novita
Kimi K2.5 is a large open-weight MoE model with 1.1 trillion parameters, known for strong vision capabilities via MoonViT-3D and agent swarm orchestration.
Tool CallingVisionLong Context
moonshotai/Kimi-K2.6
novita
Kimi-K2.6 is an open-source native multimodal agentic model developed by Moonshot AI, supporting long-horizon coding capabilities and agentic task orchestration scaling to 300 sub-agents.
Tool CallingVisionLong Context
GLM-5 is a 754B-parameter Mixture-of-Experts model from Z.ai. It features DeepSeek Sparse Attention (DSA) for efficient long-context handling, leading open-source models in long-horizon planning and agentic tasks.
Tool CallingLong Context
deepseek-ai/DeepSeek-V4-Pro
novita
DeepSeek-V4-Pro is a flagship Mixture-of-Experts (MoE) language model with 862B total parameters (49B activated) and a 1-million-token context window. It features a Hybrid Attention Architecture combining Compressed Sparse Attention and Heavily Compressed Attention for efficient long-context processing. With multiple distinct reasoning modes (such as Think Max for maximum reasoning effort), it is built for advanced mathematical reasoning, software engineering, tool use scenarios, and complex long-horizon agentic workflows.
Tool CallingLong Context
moonshotai/Kimi-K2.7-Code
novita
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
Tool CallingVisionLong Context
GPT-5.4 Mini is a compact, cost-efficient model designed for reliable performance across high-volume, everyday AI workloads. It offers fast, efficient reasoning, strong multimodal capabilities, and approaches the performance of the larger GPT-5.4 model on coding and agentic tasks while running more than 2x faster than its predecessor.
Tool CallingVisionLong Context
string
Tool CallingLong Context
GLM-5.2 is Z.ai's latest flagship open-weights MoE model engineered specifically to dominate long-horizon autonomous coding and engineering tasks. It features a solid 1-million-token context window, multiple thinking effort levels, and an improved IndexShare architecture that significantly boosts reasoning and agentic performance.
Tool CallingLong Context
Claude Sonnet 5 is the most agentic Sonnet model yet, released by Anthropic on June 30, 2026. It offers a 1M token context window and 128k max output tokens. It is designed for complex, multi-step agentic tasks like coding, tool use, and autonomous planning, achieving performance close to the more expensive Opus 4.8 model. It features adaptive thinking enabled by default and a new tokenizer that produces approximately 30% more tokens for the same text compared to Sonnet 4.6.
Tool CallingVisionLong Context
GPT-5.6 Luna is the fastest, most cost-efficient model in the GPT-5.6 family. It is designed for high-volume, latency-sensitive workloads such as classification, data extraction, request routing, and draft generation. It shares the same 1.05M context window as its siblings
Tool CallingVisionLong Context
GPT-4.1 is OpenAI's 1.6 trillion parameter multimodal reasoning system featuring a 256-expert MoE architecture with enhanced tool integration and world knowledge grounding. Optimized for enterprise applications with improved factual accuracy and reduced hallucination.
Tool CallingVisionLong Context
A large, multimodal language model with advanced reasoning capabilities and support for tools. It supports the new OpenAI Responses API with features like reasoning effort control.
Tool CallingVisionLong Context
GPT-4o (Omnimodal) is OpenAI's 1.2 trillion parameter multimodal foundation model featuring unified input processing across text, vision, and audio. Optimized for real-time interaction with enhanced reasoning and cross-modal understanding capabilities.
Tool CallingVisionLong Context
The standard 'Thinking' variant of OpenAI's GPT-5.2 model, released Dec 11, 2025. Excels at complex reasoning, coding, and multi-step projects. Optimized for professional knowledge work and serves as the balanced option between speed and capability in the GPT-5.2 family.
Tool CallingVisionLong Context
GPT-5.4 is OpenAI's flagship frontier model built for sustained, multi-step reasoning with reliable follow-through. It excels at agentic workflows, research, document analysis, and powering complex internal tools with robust instruction following.
Tool CallingVisionLong Context
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, the balanced workhorse tier of the GPT-5.6 family. Delivers highly reliable performance for everyday professional workflows, code generation, and general-purpose agentic tasks at a lower cost than Sol.
Tool CallingVisionLong Context
claude-sonnet-4-6
anthropic
Claude Sonnet 4.6 is Anthropic's balanced production model, offering near-Opus level intelligence at a much lower cost. It features a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work, and design, along with an adaptive thinking capability.
Tool CallingVisionLong Context
Kimi’s most capable flagship model to date, with 2.8 trillion parameters. Built for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. The world’s first open-source model in the 3-trillion-parameter class.
Tool CallingVisionLong Context
Claude Opus 4.6 brings Anthropic's advanced reasoning capabilities to high-stakes workflows. With full adaptive thinking support and a 1-million-token context window, it excels at complex tasks across coding, cybersecurity, financial analysis, and large-scale enterprise agents.
Tool CallingVisionLong Context
Claude Opus 4.7 excels in autonomous multi-step engineering and long-horizon agentic workflows. It features an Effort Control layer for self-correction and advanced visual intelligence, making it an ideal choice for software engineering, legal, and financial analyses.
Tool CallingVisionLong Context
Claude Opus 4.8 is a workhorse flagship model designed for high-end reasoning and long-horizon agentic coding. It offers exceptional consistency and autonomy to sustain work on complex, long-running tasks.
Tool CallingVisionLong Context
Claude Opus 5 is Anthropic's flagship model designed for complex agentic coding and enterprise work, delivering intelligence close to Claude Fable 5 at half the price. It features a 1M token context window, adaptive thinking, and state-of-the-art performance on coding and knowledge work evaluations. It is optimized for bounded, complex tasks
Tool CallingVisionLong Context
GPT-5.5 is a highly capable and efficient model featuring dynamic routing between 'Instant' and 'Thinking' modes. It provides improvements in conceptual clarity, scientific research ability, accuracy during knowledge work, and polished conversational tones.
Tool CallingVisionLong Context
GPT-5.6 Sol is the flagship model in the GPT-5.6 family, designed for the most complex reasoning, agentic coding, and deep research tasks. It features a 1.05M token context window and excels in long-horizon, multi-step workflows. It supports advanced features like programmatic tool calling, multi-agent orchestration, and a 'max' reasoning effort
Tool CallingVisionLong Context
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins.
Tool CallingVisionLong Context
OpenAI's flagship GPT-5 Pro model, released October 6, 2025. A highly capable multimodal model featuring advanced reasoning, native tool use, and strong performance on complex professional and coding tasks. It introduced a structured 'reasoning effort' parameter for controlling response depth.
Tool CallingVisionLong Context
GPT-5.4 Pro offers deeper, higher-reliability reasoning for complex production scenarios. It is engineered for high-stakes agentic workflows, long-form analysis and synthesis, complex planning, and advanced reasoning tasks where accuracy is paramount.
Tool CallingVisionLong Context
GPT-5.5 Pro is the highest-capability model in the 5.5 lineup, optimized for the hardest tasks and long-running workflows. It uses significant compute to 'think' harder before answering, making it ideal for deep research, highly complex coding, and extensive multi-step problem solving.
Tool CallingVisionLong Context
claude-3-haiku-20240307
6ce62836a581464f8ea1e6ee1d397a99
Claude 3 Haiku is Anthropic's fastest and most compact model in the Claude 3 family, optimized for speed and cost-effectiveness while maintaining strong performance for everyday tasks and high-volume operations.
Tool CallingVisionLong Context
claude-haiku-4-5-20251001
1be62feac9264117aee04eb7cdc868f2
Claude Haiku 4.5 is Anthropic's fastest and most efficient model in the Claude 4 family, optimized for speed and cost-effectiveness while maintaining high quality performance for everyday tasks.
Tool CallingVision
claude-opus-4-1-20250805
f0f29ca2cd6c408c817f31623e97e73f
Claude Opus 4.1 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning and analysis tasks. Part of the earlier Claude 4 generation alongside Claude Opus 4.
Tool CallingVisionLong Context
claude-opus-4-20250514
9d44f0f5167143d6a783d42ca0ba64e4
Claude Opus 4 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning, analysis, and challenging tasks requiring deep understanding.
Tool CallingVisionLong Context
claude-opus-4-5-20251101
14991ae9ac7d45298b4e7e05db98b4fc
Claude Opus 4.5 is Anthropic's most capable model in the Claude 4 family, offering superior performance for complex tasks requiring deep reasoning and analysis.
Tool CallingVision
claude-sonnet-4-20250514
91cad413eb7148658649629510425980
Claude Sonnet 4 is a balanced model in the Claude 4 family, offering a strong combination of intelligence, speed, and efficiency for a wide range of tasks including reasoning, coding, and analysis.
Tool CallingVisionLong Context
claude-sonnet-4-5-20250929
a40f171ceba14d75b2bf1749904d0965
Claude Sonnet 4.5 is Anthropic's smartest model and the flagship of the Claude 4 family, offering the best balance of intelligence, speed, and efficiency for everyday use. It excels at complex reasoning, coding, and analysis tasks.
Tool CallingVision
deepseek/deepseek-r1-distill-llama-8b
24c0b83803ff4ac8b32693aee54098e8
DeepSeek-R1-Distill-Llama-8B combines DeepSeek-R1 knowledge distillation with Llama architecture. Available in FP8 (H100+ only) and FP16 quantization, delivering efficient performance with 32K context.
Tool Calling
sarvamai/sarvam-105b
sarvam
Sarvam-105B is an advanced Mixture-of-Experts (MoE) model with 10.3B active parameters, designed for superior performance across a wide range of complex tasks. It is highly optimized for complex reasoning, with particular strength in agentic tasks, mathematics, and coding.
Tool CallingLong Context