Language model APIs

Compare Text and Language Model APIs

Explore language models for chat, reasoning, coding, extraction, agents, and structured generation. Every model uses the OneInfer API surface, so teams can compare and switch without rebuilding their integration.

Current catalog

94

text models with verified catalog metadata

Pricing content refreshed September 9, 2026

Lowest comparable price

Meta: Meta Llama 3 8B Instruct

Currently the lowest-priced option in the average of available input and output prices per one million tokens group. Different billing units are never mixed.

$0.040 avg / 1M tokens

View pricing and providers

Available Text Models

Comparable models are ordered by their current normalized price.

94 models

Meta: Meta Llama 3 8B Instruct

novita

Lowest price

Meta-Llama-3-8B-Instruct is an instruction-tuned variant of Meta's 8-billion-parameter transformer, optimized for dialogue and task completion. Trained with RLHF on 15T+ tokens, it delivers strong performance in reasoning, coding, and multilingual tasks while maintaining efficiency for single-GPU deployment. Features improved instruction following and reduced hallucination compared to base models.

Tool Calling

Price

$0.040 avg / 1M tokens

Context

8K

See Meta: Meta Llama 3 8B Instruct pricing

Google: Gemma 3 12B IT

novita

Gemma 3 12B Instruct is a 12-billion-parameter multimodal model from Google, built from the same research as the Gemini models. It is instruction-tuned for dialogue and excels at text generation and image understanding tasks like question answering, summarization, and reasoning.

Tool CallingVisionLong Context

Price

$0.075 avg / 1M tokens

Context

128K

See Google: Gemma 3 12B IT pricing

OpenAI: GPT Oss 20B

thinkingmachines

GPT-OSS-20B is OpenAI's efficiency-optimized 20-billion-parameter transformer featuring architectural innovations from GPT-4. Designed for accessible deployment, it delivers 85% of GPT-OSS-120B's capability at 6× lower resource requirements. Includes sparse attention mechanisms, constitutional AI safeguards, and deterministic output options.

Tool Calling

Price

$0.080 avg / 1M tokens

Context

64K

See OpenAI: GPT Oss 20B pricing

Qwen: Qwen3 Next 80B A3B Instruct

novita

Qwen3-Next-80B-A3B-Instruct is an 80B-parameter MoE model with 3B active parameters per token, optimized for instruction following and complex reasoning tasks. Features enhanced multilingual capabilities, advanced tool use, and extended context handling with improved efficiency through expert routing.

Tool CallingLong Context

Price

$0.082 avg / 1M tokens

Context

128K

See Qwen: Qwen3 Next 80B A3B Instruct pricing

Mistral AI: Mistral Nemo Instruct 2407

novita

Mistral-Nemo-Instruct-2407 is a 12-billion-parameter instruction-tuned language model developed jointly by Mistral AI and NVIDIA. It features a 128k token context window, strong multilingual capabilities, and excels at tasks like text generation, code generation, and reasoning. It uses the efficient Tekken tokenizer and is released under the Apache 2.0 license.

Tool CallingLong Context

Price

$0.105 avg / 1M tokens

Context

128K

See Mistral AI: Mistral Nemo Instruct 2407 pricing

NVIDIA: Nemotron-3 Nano 30B A3B

novita

NVIDIA Nemotron-3 Nano 30B A3B is NVIDIA's open reasoning model featuring a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture. It consists of 23 Mamba-2 and MoE layers, plus 6 Attention layers. The model is built for building specialized AI agents, chatbots, and RAG systems, providing high compute efficiency. It natively supports configurable reasoning depth through a 'thinking budget' and acts as a general-purpose reasoning and chat model intended for English, coding languages, and several other supported languages.

Tool CallingLong Context

Price

$0.125 avg / 1M tokens

Context

256K

See NVIDIA: Nemotron-3 Nano 30B A3B pricing

DeepSeek: DeepSeek R1 Distill Qwen 14B

novita

DeepSeek-R1-Distill-Qwen-14B is a knowledge-distilled model combining DeepSeek-R1's reasoning capabilities with Qwen's multilingual strengths. Features enhanced Chinese-English performance and balanced task generalization at 14B scale with FP16 precision.

Tool Calling

Price

$0.150 avg / 1M tokens

Context

32K

See DeepSeek: DeepSeek R1 Distill Qwen 14B pricing

ZAI: GLM 5.3 Flash

zai

GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.

Tool CallingVisionLong Context

Price

$0.163 avg / 1M tokens

Context

1M

See ZAI: GLM 5.3 Flash pricing

Qwen: Qwen3 235B A22B Instruct 2507

novita

Qwen3-235B-A22B-Instruct-2507 is Alibaba Cloud's frontier mixture-of-experts model featuring 235B total parameters with 22 specialized experts and 4 active per token. Designed for superhuman reasoning and multilingual mastery, it achieves state-of-the-art performance across technical, creative, and analytical domains with 256K context handling. Optimized for distributed inference across H100 GPU clusters.

Tool CallingVisionLong Context

Price

$0.165 avg / 1M tokens

Context

256K

See Qwen: Qwen3 235B A22B Instruct 2507 pricing

ByteDance: Seed 1.6 Flash

bytedance

ByteDance Seed 1.6 Flash is an ultra-fast multimodal deep-thinking model optimized for speed and cost-efficiency without compromising capability. It supports text, image, and video inputs along with tool use and function calling across an expansive 256K token context window, making it ideal for high-speed automated workflows and media-heavy processing.

Tool CallingVisionLong Context

Price

$0.188 avg / 1M tokens

Context

256K

See ByteDance: Seed 1.6 Flash pricing

DeepSeek: DeepSeek V4 Flash

novita

DeepSeek-V4-Flash is a highly efficient Mixture-of-Experts (MoE) language model featuring 158B total parameters with only 13B activated during inference. It utilizes a Hybrid Attention Architecture (combining Compressed Sparse Attention and Heavily Compressed Attention) to natively and efficiently support a 1-million-token context window. The model offers switchable reasoning modes (Non-think, Think High, and Think Max) and excels at high-speed generation for reasoning, coding, tool-calling, and agentic tasks.

Tool CallingLong Context

Price

$0.210 avg / 1M tokens

Context

1M

See DeepSeek: DeepSeek V4 Flash pricing

OpenAI: GPT 5 Nano

openai

GPT-5 Nano is a lightweight, low-latency language model designed for cost-efficient text generation, chat, and tool-calling workloads.

Tool CallingLong Context

Price

$0.225 avg / 1M tokens

Context

400K

See OpenAI: GPT 5 Nano pricing

OpenAI: GPT 4.1 Nano

openai

GPT-4.1 Nano is a highly efficient 100B-parameter variant optimized for edge deployment and mobile devices. Features 128 experts with aggressive pruning and INT8 quantization, delivering 60% of GPT-4.1's capability while consuming 75% less power.

Tool CallingVision

Price

$0.250 avg / 1M tokens

Context

64K

See OpenAI: GPT 4.1 Nano pricing

OpenAI: GPT Oss 120B

novita

GPT-OSS-120B is OpenAI's open-source 120-billion-parameter transformer model featuring architectural innovations from GPT-4. Optimized for large-scale deployment with custom FP8 quantization, it delivers state-of-the-art reasoning capabilities while maintaining 3× better throughput than comparable models. Includes constitutional AI safeguards and deterministic output options.

Tool CallingLong Context

Price

$0.264 avg / 1M tokens

Context

128K

See OpenAI: GPT Oss 120B pricing

Google: Gemma 4 26B A4b

novita

Gemma 4 26B A4B is an efficient, open-weights Mixture-of-Experts (MoE) multimodal model from Google DeepMind. With only 3.8B active parameters out of 25.2B total, it runs nearly as fast as a 4B model while delivering performance close to the 31B dense model. Features a 256K token context window, supports interleaved text and image inputs, and excels at reasoning, coding, and agentic workflows.

Tool CallingVisionLong Context

Price

$0.265 avg / 1M tokens

Context

256K

See Google: Gemma 4 26B A4b pricing

Google: Gemma 4 31B

novita

Gemma 4 31B is a state-of-the-art, open-weights multimodal dense model from Google DeepMind. It features a 256K token context window, excels at reasoning, coding, and agentic workflows, and supports interleaved text and image inputs.

Tool CallingVisionLong Context

Price

$0.270 avg / 1M tokens

Context

256K

See Google: Gemma 4 31B pricing

Qwen: Qwen3.8 Flash

openrouter

Qwen3.8-Flash is Qwen's high-speed multimodal reasoning model designed for coding, agentic workflows, visual understanding, long-context processing, and high-concurrency applications. It natively supports a 1M-token context window and can process text, images, video, lengthy documents, and large codebases. The model combines strong reasoning and generation capabilities with efficient inference, and supports built-in tools for agentic and developer workflows.

Tool CallingVisionLong Context

Price

$0.310 avg / 1M tokens

Context

1M

See Qwen: Qwen3.8 Flash pricing

DeepSeek: DeepSeek V3.2

novita

DeepSeek-V3.2 is a general-purpose Mixture-of-Experts (MoE) large language model from DeepSeek AI that harmonizes high computational efficiency with superior reasoning and agent performance. It features 685 billion total parameters, supports a 164Ktoken context window, and is natively calibrated for tool use and function calling in complex agentic workflows.

Tool CallingLong Context

Price

$0.335 avg / 1M tokens

Context

164K

See DeepSeek: DeepSeek V3.2 pricing

OpenAI: GPT 4O Mini

openai

GPT-4o Mini is a distilled 400B-parameter version of GPT-4o optimized for cost-efficient omnimodal applications. Maintains core multimodal capabilities with 70% of GPT-4o's performance at 40% of the computational cost, featuring real-time audio/image processing optimizations.

Tool CallingVisionLong Context

Price

$0.375 avg / 1M tokens

Context

128K

See OpenAI: GPT 4O Mini pricing

Qwen: Qwen 2.5 72B Instruct

novita

Qwen-2.5-72B-Instruct is Alibaba's flagship 72B parameter model featuring enhanced reasoning, 128K context, and advanced tool-calling capabilities. Excels in multilingual tasks, technical domains, and complex instruction following with improved safety alignment.

Tool CallingLong Context

Price

$0.390 avg / 1M tokens

Context

128K

See Qwen: Qwen 2.5 72B Instruct pricing

Z.ai: GLM 4.6V

novita

GLM-4.6V is a vision-language model designed for cloud and high-performance clusters. It introduces native multimodal function calling, interleaved image-text generation, and advanced document understanding. It supports a 128K context window and achieves SoTA visual understanding among models of similar scale.

Tool CallingVisionLong Context

Price

$0.600 avg / 1M tokens

Context

128K

See Z.ai: GLM 4.6V pricing

Microsoft: Wizardlm 2 8X22b

novita

WizardLM-2 8x22B is a 141-billion-parameter Mixture of Experts (MoE) model developed by Microsoft AI. It is the flagship model of the WizardLM-2 family, designed for state-of-the-art performance in complex chat, reasoning, multilingual tasks, and agent-like capabilities. It is built upon the Mixtral-8x22B architecture and trained using a fully AI-powered synthetic training system .

Tool Calling

Price

$0.620 avg / 1M tokens

Context

64K

See Microsoft: Wizardlm 2 8X22b pricing

DeepSeek: DeepSeek V3.1

novita

DeepSeek-V3.1 is a 1.3T-parameter Mixture-of-Experts model with 236B active parameters per token. It features enhanced reasoning capabilities, extended 1M token context window, and improved multilingual support with advanced tool usage and code generation capabilities.

Tool CallingLong Context

Price

$0.635 avg / 1M tokens

Context

100K

See DeepSeek: DeepSeek V3.1 pricing

OpenAI: GPT 5.6 Luna

openai

GPT-5.6 Luna is the fastest, most cost-efficient model in the GPT-5.6 family. It is designed for high-volume, latency-sensitive workloads such as classification, data extraction, request routing, and draft generation. It shares the same 1.05M context window as its siblings

Tool CallingVisionLong Context

Price

$0.700 avg / 1M tokens

Context

1.05M

See OpenAI: GPT 5.6 Luna pricing

DeepSeek: DeepSeek V3 0324

novita

DeepSeek-V3-0324 is a 671B-parameter model with 37B active parameters per token. Features strong reasoning capabilities, 128K context window, and enhanced multilingual support with advanced tool usage and code generation.

Tool CallingLong Context

Price

$0.710 avg / 1M tokens

Context

128K

See DeepSeek: DeepSeek V3 0324 pricing

OpenAI: GPT 5.4 Nano

openai

GPT-5.4 Nano is the smallest and fastest model in the GPT-5.4 lineup, designed for ultra-low latency and low-cost API usage at high throughput. It is optimized for short-turn tasks like classification, extraction, ranking, and lightweight sub-agent work.

Tool CallingVisionLong Context

Price

$0.725 avg / 1M tokens

Context

400K

See OpenAI: GPT 5.4 Nano pricing

Qwen: Qwen3 Coder 480B A35B Instruct FP8

novita

Qwen3-Coder-480B-A35B-Instruct is a massive 480B-parameter MoE model specialized in code generation and technical problem-solving. Featuring 35B active parameters per token through expert routing, it delivers state-of-the-art coding assistance with FP8 quantization for efficient deployment.

Tool CallingLong Context

Price

$0.745 avg / 1M tokens

Context

128K

See Qwen: Qwen3 Coder 480B A35B Instruct FP8 pricing

MiniMax: MiniMax M2.7

minimax

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

Tool CallingLong Context

Price

$0.750 avg / 1M tokens

Context

200K

See MiniMax: MiniMax M2.7 pricing

MiniMax: MiniMax M2.5

minimax

MiniMax M2.5 is an open-source MoE model with 230B total parameters (10B active) and a 200K token context window, released February 2026 under a Modified MIT license. SOTA in coding, agentic tool use, search, and office work — scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp. Trained across 200K+ real-world RL environments in 10+ programming languages.Matches Claude Opus 4.6 speed while completing SWE-Bench 37% faster than M2.1.

Tool CallingLong Context

Price

$0.750 avg / 1M tokens

Context

200K

See MiniMax: MiniMax M2.5 pricing

MiniMax: MiniMax M3

novita

MiniMax M3 is an open model with sparse block-level attention designed for coding, agentic workflows, and multimodal chat. It supports a 1M token context window and operates natively across text, image, and video.

Tool CallingVisionLong Context

Price

$0.750 avg / 1M tokens

Context

1M

See MiniMax: MiniMax M3 pricing

Baidu: Ernie 4.5 VL 424B A47B Base Pt

novita

ERNIE 4.5 VL is Baidu's 424-billion parameter vision-language Mixture of Experts model with 47B active parameters per token. It features advanced multimodal understanding, reasoning across text and images, and state-of-the-art performance in complex visual-language tasks.

Tool CallingVision

Price

$0.835 avg / 1M tokens

Context

32K

See Baidu: Ernie 4.5 VL 424B A47B Base Pt pricing

Meta: Muse Glimmer 30B

together_ai

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for local autonomous agentic workflows on consumer hardware. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a built-in perception encoder. It supports interleaved text and image input, controllable reasoning strength, and is highly compatible with agentic orchestration frameworks.

Tool CallingVisionLong Context

Price

$0.925 avg / 1M tokens

Context

131K

See Meta: Muse Glimmer 30B pricing

OpenAI: GPT 4.1 Mini

openai

GPT-4.1 Mini is a distilled 400B-parameter version of GPT-4.1 optimized for cost-efficient deployment. Features the same 256-expert MoE architecture with selective expert activation, delivering 80% of GPT-4.1's capability at 30% of the inference cost.

Tool CallingVisionLong Context

Price

$1.000 avg / 1M tokens

Context

262144

See OpenAI: GPT 4.1 Mini pricing

OpenAI: GPT 5 Mini

openai

OpenAI's most compact and cost-effective GPT-5 model, released August 7, 2025. Designed for high-volume, latency-sensitive applications where the full capabilities of larger models are not required, while maintaining strong performance on common tasks.

Tool CallingLong Context

Price

$1.125 avg / 1M tokens

Context

400K

See OpenAI: GPT 5 Mini pricing

ByteDance: Seed 1.8

bytedance

ByteDance Seed 1.8 is a large language model optimized for long-context reasoning and generation with a 256K context window. It excels in complex coding, multi-turn dialogue, and multi-step agentic tasks while robustly supporting vision inputs and parallel tool calling.

Tool CallingVisionLong Context

Price

$1.125 avg / 1M tokens

Context

256K

See ByteDance: Seed 1.8 pricing

ByteDance: Seed 1.6

bytedance

A model from ByteDance featuring adaptive thinking (AdaCoT) to control reasoning length. It was developed from Seed1.5-Thinking with greater training compute and integrated VLM capabilities. Supports text and image inputs with a 256K context window

Tool CallingVisionLong Context

Price

$1.125 avg / 1M tokens

Context

256K

See ByteDance: Seed 1.6 pricing

Z.ai: GLM 4.6

novita

GLM-4.6 is a large language model with improvements over GLM-4.5, featuring a 200K context window, superior coding performance, advanced reasoning, and stronger agentic capabilities in tool use and search-based agents. It also offers refined writing and more natural role-playing.

Tool CallingLong Context

Price

$1.375 avg / 1M tokens

Context

200K

See Z.ai: GLM 4.6 pricing

Z.ai: GLM 4.7

novita

GLM-4.7 is a large language model focused on coding, agentic tasks, and complex reasoning. It introduces Interleaved Thinking, Preserved Thinking, and Turn-level Thinking for improved stability and performance in multi-turn, long-horizon tasks. It shows significant gains over its predecessor in multilingual agentic coding, terminal-based tasks, and tool use.

Tool CallingLong Context

Price

$1.400 avg / 1M tokens

Context

200K

See Z.ai: GLM 4.7 pricing

Moonshot AI: Kimi K2 Instruct

novita

Kimi-K2-Instruct is Moonshot AI's flagship 100-billion-parameter instruction-tuned model featuring a hybrid dense-MoE architecture with 200K context handling. Optimized for complex reasoning, multilingual dialogue, and long-form document understanding. Specializes in Chinese and English technical domains with enhanced safety alignment.

Tool CallingLong Context

Price

$1.435 avg / 1M tokens

Context

128K

See Moonshot AI: Kimi K2 Instruct pricing

ByteDance: Dola Seed 2.1 Turbo

bytedance

Dola-Seed 2.1 Turbo is a next-generation multimodal large model by ByteDance designed for real-world productivity, high-value production tasks in enterprise R&D, and large-scale agentic scenarios. It features comprehensive upgrades in coding engineering delivery, long-horizon agent task execution, and multimodal understanding, along with stronger autonomous planning and dynamic self-repair capabilities.

Tool CallingVisionLong Context

Price

$1.500 avg / 1M tokens

Context

256K

See ByteDance: Dola Seed 2.1 Turbo pricing

Moonshot AI: Kimi K2 0905

novita

Kimi-K2-0905 is a 72B-parameter multimodal language model optimized for long-context understanding and complex reasoning. Features enhanced multilingual capabilities, advanced tool usage, and strong performance in mathematical and logical reasoning tasks with extended context handling.

Tool CallingVisionLong Context

Price

$1.550 avg / 1M tokens

Context

128K

See Moonshot AI: Kimi K2 0905 pricing

Qwen: Qwen3 235B A22B Thinking 2507

novita

Qwen3-235B-A22B-Thinking-2507 is Alibaba Cloud's cognitive architecture model specializing in advanced reasoning, theory of mind simulation, and multi-agent collaboration. Built on the Qwen3-235B MoE foundation, it features specialized 'cognitive experts' and a neuro-symbolic execution engine for human-like problem-solving. Excels at strategic planning, counterfactual reasoning, and complex system modeling with 256K context retention.

Tool CallingLong Context

Price

$1.650 avg / 1M tokens

Context

256K

See Qwen: Qwen3 235B A22B Thinking 2507 pricing

Qwen: Qwen3.8 27B

akashml

Qwen3.8-27B is a compact, deployment-friendly dense multimodal model from Alibaba's Qwen3.8 family. Natively supporting text, image, and video inputs, it handles tasks ranging from STEM diagrams and document understanding to long-form video analysis. It features flexible thinking control with adjustable reasoning effort and is designed for autonomous planning, coding, professional work, multimodal reasoning, and long-horizon agentic workflows.

Tool CallingVisionLong Context

Price

$1.710 avg / 1M tokens

Context

256K

See Qwen: Qwen3.8 27B pricing

ByteDance: Dola Seed 2.0 Code

bytedance

Dola-Seed 2.0 Code is ByteDance's dedicated coding model optimized for enterprise-level programming needs. It excels in agentic task planning with rational decomposition and dynamic self-repair, while inheriting the Seed 2.0 family's robust multimodal perception to read charts, screenshots, and diagrams for contextually accurate code generation.

Tool CallingVisionLong Context

Price

$1.750 avg / 1M tokens

Context

256K

See ByteDance: Dola Seed 2.0 Code pricing

Moonshot AI: Kimi K2.5

novita

Kimi K2.5 is a large open-weight MoE model with 1.1 trillion parameters, known for strong vision capabilities via MoonViT-3D and agent swarm orchestration.

Tool CallingVisionLong Context

Price

$1.800 avg / 1M tokens

Context

256K

See Moonshot AI: Kimi K2.5 pricing

XAI: Grok 4.20

grok

Grok 4.20 Reasoning is a high-performance model by xAI designed for complex logic and math, scientific and technical analysis, multi-step investigations, and high-stakes tasks where accuracy matters most. It incorporates a 'thinking' mechanism before responding and offers strict prompt adherence alongside agentic tool calling.

Tool CallingVisionLong Context

Price

$1.875 avg / 1M tokens

Context

1M

See XAI: Grok 4.20 pricing

XAI: Grok 4.3

grok

Grok 4.3 is a flagship reasoning model from xAI leveraging a four-agent deliberation system where internal sub-agents debate the best response. It features native video input and leverages real-time data streams from the X platform for current-events knowledge. It is optimized for long-document analysis, conversational research, and complex multi-step agentic workflows.

Tool CallingVisionLong Context

Price

$1.875 avg / 1M tokens

Context

1M

See XAI: Grok 4.3 pricing

Moonshot AI: Kimi K2.6

novita

Kimi-K2.6 is an open-source native multimodal agentic model developed by Moonshot AI, supporting long-horizon coding capabilities and agentic task orchestration scaling to 300 sub-agents.

Tool CallingVisionLong Context

Price

$2.100 avg / 1M tokens

Context

256K

See Moonshot AI: Kimi K2.6 pricing

Z.ai: GLM 5

novita

GLM-5 is a 754B-parameter Mixture-of-Experts model from Z.ai. It features DeepSeek Sparse Attention (DSA) for efficient long-context handling, leading open-source models in long-horizon planning and agentic tasks.

Tool CallingLong Context

Price

$2.100 avg / 1M tokens

Context

200K

See Z.ai: GLM 5 pricing

NVIDIA: Nemotron-3 Ultra 550B A55B

together_ai

NVIDIA Nemotron-3 Ultra 550B A55B is a frontier-scale large language model from NVIDIA, designed for strong agentic, reasoning, and conversational capabilities. It employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2, MoE, and Attention layers, alongside Multi-Token Prediction (MTP) for faster inference. With a 1-million-token context window and a configurable reasoning mode, it excels at complex autonomous agents, long-context analysis, and deep research workflows.

Tool CallingLong Context

Price

$2.100 avg / 1M tokens

Context

1M

See NVIDIA: Nemotron-3 Ultra 550B A55B pricing

DeepSeek: DeepSeek V4 Pro

novita

DeepSeek-V4-Pro is a flagship Mixture-of-Experts (MoE) language model with 862B total parameters (49B activated) and a 1-million-token context window. It features a Hybrid Attention Architecture combining Compressed Sparse Attention and Heavily Compressed Attention for efficient long-context processing. With multiple distinct reasoning modes (such as Think Max for maximum reasoning effort), it is built for advanced mathematical reasoning, software engineering, tool use scenarios, and complex long-horizon agentic workflows.

Tool CallingLong Context

Price

$2.400 avg / 1M tokens

Context

1M

See DeepSeek: DeepSeek V4 Pro pricing

Moonshot AI: Kimi K2.7 Code

novita

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Tool CallingVisionLong Context

Price

$2.475 avg / 1M tokens

Context

256K

See Moonshot AI: Kimi K2.7 Code pricing

OpenAI: GPT 5.4 Mini

openai

GPT-5.4 Mini is a compact, cost-efficient model designed for reliable performance across high-volume, everyday AI workloads. It offers fast, efficient reasoning, strong multimodal capabilities, and approaches the performance of the larger GPT-5.4 model on coding and agentic tasks while running more than 2x faster than its predecessor.

Tool CallingVisionLong Context

Price

$2.625 avg / 1M tokens

Context

400K

See OpenAI: GPT 5.4 Mini pricing

Meta: Muse Spark 1.3

meta

Muse Spark 1.3 is Meta's natively multimodal reasoning model optimized for long-horizon agentic workflows, advanced coding, tool use, computer interaction, and complex multi-step tasks. It is designed to sustain multiple workflows within long conversations, generate and manage context across conflicting sources, identify and correct gaps in its plans, and efficiently complete real-world engineering and knowledge-work tasks. Compared with Muse Spark 1.2, it improves coding and agentic performance while using fewer tool calls and fewer tokens.

Tool CallingVisionLong Context

Price

$2.750 avg / 1M tokens

Context

1M

See Meta: Muse Spark 1.3 pricing

Zai: GLM 5.3

zai

GLM-5.3 is Z.ai's advanced Mixture-of-Experts model, building upon the GLM-5.2 base architecture with significantly scaled post-training. It is designed specifically for complex coding, long-horizon tasks, and autonomous agent workflows. The model sets a new state-of-the-art for open models in agentic coding and demonstrates strong emergent cybersecurity capabilities, excelling in vulnerability discovery and exploitation chaining.

Tool CallingLong Context

Price

$2.850 avg / 1M tokens

Context

1M

See Zai: GLM 5.3 pricing

Z.ai: GLM 5.2

novita

GLM-5.2 is Z.ai's latest flagship open-weights MoE model engineered specifically to dominate long-horizon autonomous coding and engineering tasks. It features a solid 1-million-token context window, multiple thinking effort levels, and an improved IndexShare architecture that significantly boosts reasoning and agentic performance.

Tool CallingLong Context

Price

$2.900 avg / 1M tokens

Context

1M

See Z.ai: GLM 5.2 pricing

Anthropic: Claude Sonnet 5

anthropic

Claude Sonnet 5 is the most agentic Sonnet model yet, released by Anthropic on June 30, 2026. It offers a 1M token context window and 128k max output tokens. It is designed for complex, multi-step agentic tasks like coding, tool use, and autonomous planning, achieving performance close to the more expensive Opus 4.8 model. It features adaptive thinking enabled by default and a new tokenizer that produces approximately 30% more tokens for the same text compared to Sonnet 4.6.

Tool CallingVisionLong Context

Price

$3.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Sonnet 5 pricing

Qwen: Qwen3.8 Max

novita

Qwen3.8-Max is Alibaba's 2.4-trillion-parameter flagship multimodal large language model. Designed for long-horizon professional workflows, it excels in full-stack coding, agentic development, data analysis, and office automation. It processes text, images, video, and documents natively, and is positioned as a leading frontier model.

Tool CallingVisionLong Context

Price

$4.000 avg / 1M tokens

Context

1M

See Qwen: Qwen3.8 Max pricing

XAI: Grok 4.5

grok

Grok 4.5 is xAI's frontier model built for coding, agentic tasks, and knowledge work. Trained on datasets spanning science, engineering, math, and code (in partnership with Cursor), it features a 1.5 trillion parameter architecture and exceeds comparable models in real-world software engineering tasks.

Tool CallingVisionLong Context

Price

$4.000 avg / 1M tokens

Context

500K

See XAI: Grok 4.5 pricing

XAI: Grok 4.6

grok

Grok 4.6 is an incremental flagship upgrade over Grok 4.5, utilizing the same 1.5 trillion parameter foundation but enhanced with superior supervised fine-tuning and reinforcement learning via the Grok Build harness. It brings frontier intelligence to long-running AI agents, complex multi-step tasks, and software development, offering aggressive verification and highly competitive cost-per-task efficiency.

Tool CallingVisionLong Context

Price

$4.000 avg / 1M tokens

Context

500K

See XAI: Grok 4.6 pricing

Thinking Machines: Inkling Small

thinkingmachines

Inkling-Small is a general-purpose multimodal Mixture-of-Experts (MoE) model featuring 266B total parameters with only 12B active during inference. Natively processing text, image, and audio inputs, it supports a 256K token context window and variable thinking effort. Despite being a quarter of the size of the flagship Inkling model, it matches or exceeds its predecessor on key coding and agentic tasks while requiring significantly less compute.

Tool CallingVisionLong Context

Price

$4.050 avg / 1M tokens

Context

256K

See Thinking Machines: Inkling Small pricing

OpenAI: GPT 4.1

openai

GPT-4.1 is OpenAI's 1.6 trillion parameter multimodal reasoning system featuring a 256-expert MoE architecture with enhanced tool integration and world knowledge grounding. Optimized for enterprise applications with improved factual accuracy and reduced hallucination.

Tool CallingVisionLong Context

Price

$5.000 avg / 1M tokens

Context

524288

See OpenAI: GPT 4.1 pricing

OpenAI: GPT 5

openai

A large, multimodal language model with advanced reasoning capabilities and support for tools. It supports the new OpenAI Responses API with features like reasoning effort control.

Tool CallingVisionLong Context

Price

$5.625 avg / 1M tokens

Context

400K

See OpenAI: GPT 5 pricing

OpenAI: GPT 5.1

openai

OpenAI's GPT-5.1 model, released November 13, 2025. Features a strong multimodal and reasoning capabilities, and is optimized for complex agentic workflows. It introduced significant improvements in coding, reasoning, and tool use over previous generations.

Tool CallingVisionLong Context

Price

$5.625 avg / 1M tokens

Context

400K

See OpenAI: GPT 5.1 pricing

OpenAI: GPT 4O

openai

GPT-4o (Omnimodal) is OpenAI's 1.2 trillion parameter multimodal foundation model featuring unified input processing across text, vision, and audio. Optimized for real-time interaction with enhanced reasoning and cross-modal understanding capabilities.

Tool CallingVisionLong Context

Price

$6.250 avg / 1M tokens

Context

128K

See OpenAI: GPT 4O pricing

OpenAI: GPT 5.6 Terra

openai

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, the balanced workhorse tier of the GPT-5.6 family. Delivers highly reliable performance for everyday professional workflows, code generation, and general-purpose agentic tasks at a lower cost than Sol.

Tool CallingVisionLong Context

Price

$7.000 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.6 Terra pricing

OpenAI: GPT 5.2

openai

The standard 'Thinking' variant of OpenAI's GPT-5.2 model, released Dec 11, 2025. Excels at complex reasoning, coding, and multi-step projects. Optimized for professional knowledge work and serves as the balanced option between speed and capability in the GPT-5.2 family.

Tool CallingVisionLong Context

Price

$7.875 avg / 1M tokens

Context

400K

See OpenAI: GPT 5.2 pricing

OpenAI: GPT 5.4

openai

GPT-5.4 is OpenAI's flagship frontier model built for sustained, multi-step reasoning with reliable follow-through. It excels at agentic workflows, research, document analysis, and powering complex internal tools with robust instruction following.

Tool CallingVisionLong Context

Price

$8.750 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.4 pricing

Anthropic: Claude Sonnet 4 6

anthropic

Claude Sonnet 4.6 is Anthropic's balanced production model, offering near-Opus level intelligence at a much lower cost. It features a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work, and design, along with an adaptive thinking capability.

Tool CallingVisionLong Context

Price

$9.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Sonnet 4 6 pricing

Moonshot AI: Kimi K3

novita

Kimi’s most capable flagship model to date, with 2.8 trillion parameters. Built for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. The world’s first open-source model in the 3-trillion-parameter class.

Tool CallingVisionLong Context

Price

$9.000 avg / 1M tokens

Context

1M

See Moonshot AI: Kimi K3 pricing

OpenAI: GPT 5.6 Sol

openai

GPT-5.6 Sol is the flagship model in the GPT-5.6 family, designed for the most complex reasoning, agentic coding, and deep research tasks. It features a 1.05M token context window and excels in long-horizon, multi-step workflows. It supports advanced features like programmatic tool calling, multi-agent orchestration, and a 'max' reasoning effort

Tool CallingVisionLong Context

Price

$12.000 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.6 Sol pricing

Thinking Machines: Inkling

thinkingmachines

Inkling is a flagship open-weights multimodal Mixture-of-Experts (MoE) model from Thinking Machines Lab featuring 975B total parameters with 41B active per token. It natively processes text, image, and audio inputs and supports a 256K context window. The model offers a controllable 'thinking-effort' dial to balance compute cost against reasoning quality, excelling across coding, agentic workflows, and general reasoning tasks.

Tool CallingVisionLong Context

Price

$13.100 avg / 1M tokens

Context

256K

See Thinking Machines: Inkling pricing

Anthropic: Claude Opus 4 6

anthropic

Claude Opus 4.6 brings Anthropic's advanced reasoning capabilities to high-stakes workflows. With full adaptive thinking support and a 1-million-token context window, it excels at complex tasks across coding, cybersecurity, financial analysis, and large-scale enterprise agents.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Opus 4 6 pricing

Anthropic: Claude Opus 4 7

anthropic

Claude Opus 4.7 excels in autonomous multi-step engineering and long-horizon agentic workflows. It features an Effort Control layer for self-correction and advanced visual intelligence, making it an ideal choice for software engineering, legal, and financial analyses.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Opus 4 7 pricing

Anthropic: Claude Opus 4 8

anthropic

Claude Opus 4.8 is a workhorse flagship model designed for high-end reasoning and long-horizon agentic coding. It offers exceptional consistency and autonomy to sustain work on complex, long-running tasks.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Opus 4 8 pricing

Anthropic: Claude Opus 5

anthropic

Claude Opus 5 is Anthropic's flagship model designed for complex agentic coding and enterprise work, delivering intelligence close to Claude Fable 5 at half the price. It features a 1M token context window, adaptive thinking, and state-of-the-art performance on coding and knowledge work evaluations. It is optimized for bounded, complex tasks

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Opus 5 pricing

OpenAI: GPT 5.5

openai

GPT-5.5 is a highly capable and efficient model featuring dynamic routing between 'Instant' and 'Thinking' modes. It provides improvements in conceptual clarity, scientific research ability, accuracy during knowledge work, and polished conversational tones.

Tool CallingVisionLong Context

Price

$17.500 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.5 pricing

Anthropic: Claude Fable 5

anthropic

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins.

Tool CallingVisionLong Context

Price

$30.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Fable 5 pricing

Anthropic: Claude Fable 5.1

anthropic

Claude Fable 5.1 is Anthropic's frontier intelligence model designed for demanding reasoning, long-horizon agentic work, advanced software engineering, multistep research, and complex enterprise workflows. It is built for tasks that can span hours and multiple applications, with strong planning, tool orchestration, failure recovery, coding, document understanding, and vision capabilities. The model supports a 1M-token context window, adaptive reasoning that is always enabled, and up to 128K output tokens.

Tool CallingVisionLong Context

Price

$30.000 avg / 1M tokens

Context

1M

See Anthropic: Claude Fable 5.1 pricing

OpenAI: GPT 6 Astra

openai

GPT-6 Astra is OpenAI's flagship model, released 3 September 2026. It costs $10 per million input tokens and $50 per million output, with a 1,050,000-token context window and 128,000 max output tokens. It accepts text and images, and supports streaming, function calling and structured outputs.

Tool CallingVisionLong Context

Price

$30.000 avg / 1M tokens

Context

1.05M

See OpenAI: GPT 6 Astra pricing

OpenAI: GPT 5 Pro

openai

OpenAI's flagship GPT-5 Pro model, released October 6, 2025. A highly capable multimodal model featuring advanced reasoning, native tool use, and strong performance on complex professional and coding tasks. It introduced a structured 'reasoning effort' parameter for controlling response depth.

Tool CallingVisionLong Context

Price

$67.500 avg / 1M tokens

Context

400K

See OpenAI: GPT 5 Pro pricing

OpenAI: GPT 5.2 Pro

openai

OpenAI's flagship GPT-5.2 Pro model, Designed for professional knowledge work and agentic workflows, exclusive 'xhigh' reasoning effort, and top-tier performance on complex reasoning, coding, and scientific benchmarks.

Tool CallingVisionLong Context

Price

$94.500 avg / 1M tokens

Context

400K

See OpenAI: GPT 5.2 Pro pricing

OpenAI: GPT 5.4 Pro

openai

GPT-5.4 Pro offers deeper, higher-reliability reasoning for complex production scenarios. It is engineered for high-stakes agentic workflows, long-form analysis and synthesis, complex planning, and advanced reasoning tasks where accuracy is paramount.

Tool CallingVisionLong Context

Price

$105.000 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.4 Pro pricing

OpenAI: GPT 5.5 Pro

openai

GPT-5.5 Pro is the highest-capability model in the 5.5 lineup, optimized for the hardest tasks and long-running workflows. It uses significant compute to 'think' harder before answering, making it ideal for deep research, highly complex coding, and extensive multi-step problem solving.

Tool CallingVisionLong Context

Price

$105.000 avg / 1M tokens

Context

1M

See OpenAI: GPT 5.5 Pro pricing

Anthropic: Claude 3 Haiku 20240307

6ce62836a581464f8ea1e6ee1d397a99

Claude 3 Haiku is Anthropic's fastest and most compact model in the Claude 3 family, optimized for speed and cost-effectiveness while maintaining strong performance for everyday tasks and high-volume operations.

Tool CallingVisionLong Context

Anthropic: Claude Haiku 4 5 20251001

1be62feac9264117aee04eb7cdc868f2

Claude Haiku 4.5 is Anthropic's fastest and most efficient model in the Claude 4 family, optimized for speed and cost-effectiveness while maintaining high quality performance for everyday tasks.

Tool CallingVision

Anthropic: Claude Opus 4 1 20250805

f0f29ca2cd6c408c817f31623e97e73f

Claude Opus 4.1 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning and analysis tasks. Part of the earlier Claude 4 generation alongside Claude Opus 4.

Tool CallingVisionLong Context

Anthropic: Claude Opus 4 20250514

9d44f0f5167143d6a783d42ca0ba64e4

Claude Opus 4 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning, analysis, and challenging tasks requiring deep understanding.

Tool CallingVisionLong Context

Anthropic: Claude Opus 4 5 20251101

14991ae9ac7d45298b4e7e05db98b4fc

Claude Opus 4.5 is Anthropic's most capable model in the Claude 4 family, offering superior performance for complex tasks requiring deep reasoning and analysis.

Tool CallingVision

Anthropic: Claude Sonnet 4 20250514

91cad413eb7148658649629510425980

Claude Sonnet 4 is a balanced model in the Claude 4 family, offering a strong combination of intelligence, speed, and efficiency for a wide range of tasks including reasoning, coding, and analysis.

Tool CallingVisionLong Context

Anthropic: Claude Sonnet 4 5 20250929

a40f171ceba14d75b2bf1749904d0965

Claude Sonnet 4.5 is Anthropic's smartest model and the flagship of the Claude 4 family, offering the best balance of intelligence, speed, and efficiency for everyday use. It excels at complex reasoning, coding, and analysis tasks.

Tool CallingVision

DeepSeek: DeepSeek R1 Distill Llama 8B

24c0b83803ff4ac8b32693aee54098e8

DeepSeek-R1-Distill-Llama-8B combines DeepSeek-R1 knowledge distillation with Llama architecture. Available in FP8 (H100+ only) and FP16 quantization, delivering efficient performance with 32K context.

Tool Calling

Sarvam AI: Sarvam 105B

sarvam

Sarvam-105B is an advanced Mixture-of-Experts (MoE) model with 10.3B active parameters, designed for superior performance across a wide range of complex tasks. It is highly optimized for complex reasoning, with particular strength in agentic tasks, mathematics, and coding.

Tool CallingLong Context

Price

See providers

Context

128K

See Sarvam AI: Sarvam 105B pricing

How to choose

Conversational applications
AI agents
Code generation
RAG and document analysis

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How are text model prices compared?

When input and output token prices are available, OneInfer uses their average as a simple comparison score. The full input and output rates remain visible on each model page.

Can I switch text models without changing my integration?

Yes. Models are accessed through the unified OneInfer API; change the model identifier while keeping the same authentication and request pattern.