Language model APIs

Compare Text and Language Model APIs

Explore language models for chat, reasoning, coding, extraction, agents, and structured generation. Every model uses the OneInfer API surface, so teams can compare and switch without rebuilding their integration.

Current catalog

75

text models with verified catalog metadata

Pricing content refreshed July 25, 2026

Lowest comparable price

meta-llama/Meta-Llama-3-8B-Instruct

Currently the lowest-priced option in the average of available input and output prices per one million tokens group. Different billing units are never mixed.

$0.040 avg / 1M tokens

View pricing and providers

Available Text Models

Comparable models are ordered by their current normalized price.

75 models

meta-llama/Meta-Llama-3-8B-Instruct

novita

Lowest price

Meta-Llama-3-8B-Instruct is an instruction-tuned variant of Meta's 8-billion-parameter transformer, optimized for dialogue and task completion. Trained with RLHF on 15T+ tokens, it delivers strong performance in reasoning, coding, and multilingual tasks while maintaining efficiency for single-GPU deployment. Features improved instruction following and reduced hallucination compared to base models.

Tool Calling

Price

$0.040 avg / 1M tokens

Context

8K

See meta-llama/Meta-Llama-3-8B-Instruct pricing

google/gemma-3-12b-it

novita

Gemma 3 12B Instruct is a 12-billion-parameter multimodal model from Google, built from the same research as the Gemini models. It is instruction-tuned for dialogue and excels at text generation and image understanding tasks like question answering, summarization, and reasoning.

Tool CallingVisionLong Context

Price

$0.075 avg / 1M tokens

Context

128K

See google/gemma-3-12b-it pricing

Qwen/Qwen3-Next-80B-A3B-Instruct

novita

Qwen3-Next-80B-A3B-Instruct is an 80B-parameter MoE model with 3B active parameters per token, optimized for instruction following and complex reasoning tasks. Features enhanced multilingual capabilities, advanced tool use, and extended context handling with improved efficiency through expert routing.

Tool CallingLong Context

Price

$0.082 avg / 1M tokens

Context

128K

See Qwen/Qwen3-Next-80B-A3B-Instruct pricing

mistralai/Mistral-Nemo-Instruct-2407

novita

Mistral-Nemo-Instruct-2407 is a 12-billion-parameter instruction-tuned language model developed jointly by Mistral AI and NVIDIA. It features a 128k token context window, strong multilingual capabilities, and excels at tasks like text generation, code generation, and reasoning. It uses the efficient Tekken tokenizer and is released under the Apache 2.0 license.

Tool CallingLong Context

Price

$0.105 avg / 1M tokens

Context

128K

See mistralai/Mistral-Nemo-Instruct-2407 pricing

openai/gpt-oss-20b

openai

GPT-OSS-20B is OpenAI's efficiency-optimized 20-billion-parameter transformer featuring architectural innovations from GPT-4. Designed for accessible deployment, it delivers 85% of GPT-OSS-120B's capability at 6× lower resource requirements. Includes sparse attention mechanisms, constitutional AI safeguards, and deterministic output options.

Tool Calling

Price

$0.125 avg / 1M tokens

Context

64K

See openai/gpt-oss-20b pricing

deepseek/deepseek-r1-distill-qwen-14b

novita

DeepSeek-R1-Distill-Qwen-14B is a knowledge-distilled model combining DeepSeek-R1's reasoning capabilities with Qwen's multilingual strengths. Features enhanced Chinese-English performance and balanced task generalization at 14B scale with FP16 precision.

Tool Calling

Price

$0.150 avg / 1M tokens

Context

32K

See deepseek/deepseek-r1-distill-qwen-14b pricing

Qwen/Qwen3-235B-A22B-Instruct-2507

novita

Qwen3-235B-A22B-Instruct-2507 is Alibaba Cloud's frontier mixture-of-experts model featuring 235B total parameters with 22 specialized experts and 4 active per token. Designed for superhuman reasoning and multilingual mastery, it achieves state-of-the-art performance across technical, creative, and analytical domains with 256K context handling. Optimized for distributed inference across H100 GPU clusters.

Tool CallingVisionLong Context

Price

$0.165 avg / 1M tokens

Context

256K

See Qwen/Qwen3-235B-A22B-Instruct-2507 pricing

gemma2-9b-it

groq

Gemma2 9B-IT is Google's lightweight 9B-parameter instruction-tuned model optimized for responsible deployment. Features enhanced reasoning, multilingual support, and hardware-efficient design for accessible AI applications with built-in safety filters and tool integration capabilities.

Tool Calling

Price

$0.200 avg / 1M tokens

Context

8K

See gemma2-9b-it pricing

google/gemma-2-9b-it

groq

Gemma2 9B-IT is Google's lightweight 9.24B-parameter instruction-tuned model optimized for responsible deployment. Features enhanced reasoning, multilingual support, and hardware-efficient design for accessible AI applications with built-in safety filters and tool integration capabilities.

Tool Calling

Price

$0.200 avg / 1M tokens

Context

8K

See google/gemma-2-9b-it pricing

meta-llama/llama-4-scout-17b-16e-instruct

groq

Llama-4-Scout-17B-16E is Meta's efficiency-optimized MoE model featuring 16 experts with ~2B active parameters per token. Designed for exploratory tasks with enhanced web navigation, research synthesis, and information discovery capabilities at low resource cost.

Tool CallingLong Context

Price

$0.225 avg / 1M tokens

Context

128K

See meta-llama/llama-4-scout-17b-16e-instruct pricing

gpt-5-nano

openai

GPT-5 Nano is a lightweight, low-latency language model designed for cost-efficient text generation, chat, and tool-calling workloads.

Tool CallingLong Context

Price

$0.225 avg / 1M tokens

Context

400K

See gpt-5-nano pricing

gpt-4.1-nano

openai

GPT-4.1 Nano is a highly efficient 100B-parameter variant optimized for edge deployment and mobile devices. Features 128 experts with aggressive pruning and INT8 quantization, delivering 60% of GPT-4.1's capability while consuming 75% less power.

Tool CallingVision

Price

$0.250 avg / 1M tokens

Context

64K

See gpt-4.1-nano pricing

meta-llama/llama-3.3-70b-instruct

groq

Llama-3.3-70B-Instruct is Meta's flagship 70B parameter model featuring enhanced reasoning, 128K context, and enterprise-grade instruction following. Represents a significant evolution over Llama 3.1 with improved tool integration, safety alignment, and complex task handling.

Tool CallingLong Context

Price

$0.260 avg / 1M tokens

Context

128K

See meta-llama/llama-3.3-70b-instruct pricing

google/gemma-4-26B-A4B

novita

Gemma 4 26B A4B is an efficient, open-weights Mixture-of-Experts (MoE) multimodal model from Google DeepMind. With only 3.8B active parameters out of 25.2B total, it runs nearly as fast as a 4B model while delivering performance close to the 31B dense model. Features a 256K token context window, supports interleaved text and image inputs, and excels at reasoning, coding, and agentic workflows.

Tool CallingVisionLong Context

Price

$0.265 avg / 1M tokens

Context

256K

See google/gemma-4-26B-A4B pricing

google/gemma-4-31B

novita

Gemma 4 31B is a state-of-the-art, open-weights multimodal dense model from Google DeepMind. It features a 256K token context window, excels at reasoning, coding, and agentic workflows, and supports interleaved text and image inputs.

Tool CallingVisionLong Context

Price

$0.270 avg / 1M tokens

Context

256K

See google/gemma-4-31B pricing

qwen/qwen3-30b-a3b-fp8

novita

Qwen3-30B-A3B-FP8 is a 30B-parameter MoE model with 3B active parameters per token, optimized with FP8 quantization. Balances high performance with practical deployment requirements across reasoning, multilingual, and coding tasks.

Tool CallingLong Context

Price

$0.275 avg / 1M tokens

Context

128K

See qwen/qwen3-30b-a3b-fp8 pricing

openai/gpt-oss-120b

novita

GPT-OSS-120B is OpenAI's open-source 120-billion-parameter transformer model featuring architectural innovations from GPT-4. Optimized for large-scale deployment with custom FP8 quantization, it delivers state-of-the-art reasoning capabilities while maintaining 3× better throughput than comparable models. Includes constitutional AI safeguards and deterministic output options.

Tool CallingLong Context

Price

$0.300 avg / 1M tokens

Context

128K

See openai/gpt-oss-120b pricing

gpt-4o-mini

openai

GPT-4o Mini is a distilled 400B-parameter version of GPT-4o optimized for cost-efficient omnimodal applications. Maintains core multimodal capabilities with 70% of GPT-4o's performance at 40% of the computational cost, featuring real-time audio/image processing optimizations.

Tool CallingVisionLong Context

Price

$0.375 avg / 1M tokens

Context

128K

See gpt-4o-mini pricing

qwen/qwen-2.5-72b-instruct

novita

Qwen-2.5-72B-Instruct is Alibaba's flagship 72B parameter model featuring enhanced reasoning, 128K context, and advanced tool-calling capabilities. Excels in multilingual tasks, technical domains, and complex instruction following with improved safety alignment.

Tool CallingLong Context

Price

$0.390 avg / 1M tokens

Context

128K

See qwen/qwen-2.5-72b-instruct pricing

meta-llama/llama-4-maverick-17b-128e-instruct

groq

Llama 4 Maverick 17B-128E Instruct is Meta's 17B-parameter MoE instruction model featuring 128 experts with 4 active per token. Optimized for complex task decomposition and tool-integrated reasoning with constitutional AI safety, delivering expert-level performance at accessible computational requirements.

Tool CallingLong Context

deepseek-chat

deepseek

DeepSeek-Chat is a 67B-parameter conversational AI optimized for helpful, honest, and engaging dialogue. Features instruction-following capabilities, emotional intelligence, and context-aware responses with safety alignment across 30+ conversation domains.

Tool CallingLong Context

Price

$0.585 avg / 1M tokens

Context

128K

See deepseek-chat pricing

microsoft/WizardLM-2-8x22B

novita

WizardLM-2 8x22B is a 141-billion-parameter Mixture of Experts (MoE) model developed by Microsoft AI. It is the flagship model of the WizardLM-2 family, designed for state-of-the-art performance in complex chat, reasoning, multilingual tasks, and agent-like capabilities. It is built upon the Mixtral-8x22B architecture and trained using a fully AI-powered synthetic training system .

Tool Calling

Price

$0.620 avg / 1M tokens

Context

64K

See microsoft/WizardLM-2-8x22B pricing

deepseek-ai/DeepSeek-V3.1

novita

DeepSeek-V3.1 is a 1.3T-parameter Mixture-of-Experts model with 236B active parameters per token. It features enhanced reasoning capabilities, extended 1M token context window, and improved multilingual support with advanced tool usage and code generation capabilities.

Tool CallingLong Context

Price

$0.635 avg / 1M tokens

Context

100K

See deepseek-ai/DeepSeek-V3.1 pricing

deepseek-ai/DeepSeek-V3-0324

novita

DeepSeek-V3-0324 is a 671B-parameter model with 37B active parameters per token. Features strong reasoning capabilities, 128K context window, and enhanced multilingual support with advanced tool usage and code generation.

Tool CallingLong Context

Price

$0.710 avg / 1M tokens

Context

128K

See deepseek-ai/DeepSeek-V3-0324 pricing

gpt-5.4-nano

openai

GPT-5.4 Nano is the smallest and fastest model in the GPT-5.4 lineup, designed for ultra-low latency and low-cost API usage at high throughput. It is optimized for short-turn tasks like classification, extraction, ranking, and lightweight sub-agent work.

Tool CallingVisionLong Context

Price

$0.725 avg / 1M tokens

Context

400K

See gpt-5.4-nano pricing

Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8

novita

Qwen3-Coder-480B-A35B-Instruct is a massive 480B-parameter MoE model specialized in code generation and technical problem-solving. Featuring 35B active parameters per token through expert routing, it delivers state-of-the-art coding assistance with FP8 quantization for efficient deployment.

Tool CallingLong Context

Price

$0.745 avg / 1M tokens

Context

128K

See Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 pricing

MiniMax-M2.7

minimax

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

Tool CallingLong Context

Price

$0.750 avg / 1M tokens

Context

200K

See MiniMax-M2.7 pricing

MiniMax-M2.5

minimax

MiniMax M2.5 is an open-source MoE model with 230B total parameters (10B active) and a 200K token context window, released February 2026 under a Modified MIT license. SOTA in coding, agentic tool use, search, and office work — scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp. Trained across 200K+ real-world RL environments in 10+ programming languages.Matches Claude Opus 4.6 speed while completing SWE-Bench 37% faster than M2.1.

Tool CallingLong Context

Price

$0.750 avg / 1M tokens

Context

200K

See MiniMax-M2.5 pricing

MiniMaxAI/MiniMax-M3

novita

MiniMax M3 is an open model with sparse block-level attention designed for coding, agentic workflows, and multimodal chat. It supports a 1M token context window and operates natively across text, image, and video.

Tool CallingVisionLong Context

Price

$0.750 avg / 1M tokens

Context

1M

See MiniMaxAI/MiniMax-M3 pricing

baidu/ERNIE-4.5-VL-424B-A47B-Base-PT

novita

ERNIE 4.5 VL is Baidu's 424-billion parameter vision-language Mixture of Experts model with 47B active parameters per token. It features advanced multimodal understanding, reasoning across text and images, and state-of-the-art performance in complex visual-language tasks.

Tool CallingVision

Price

$0.835 avg / 1M tokens

Context

32K

See baidu/ERNIE-4.5-VL-424B-A47B-Base-PT pricing

deepseek-r1-distill-llama-70b

groq

DeepSeek-R1-Distill-Llama-70B is a hybrid 70B-parameter model combining DeepSeek's reasoning capabilities with Llama-70B's architectural efficiency. Knowledge-distilled from DeepSeek-R1 to Llama-3 architecture, delivering 92% of R1's performance with 60% lower inference cost and enhanced tool integration.

Tool CallingLong Context

Price

$0.870 avg / 1M tokens

Context

128K

See deepseek-r1-distill-llama-70b pricing

gpt-4.1-mini

openai

GPT-4.1 Mini is a distilled 400B-parameter version of GPT-4.1 optimized for cost-efficient deployment. Features the same 256-expert MoE architecture with selective expert activation, delivering 80% of GPT-4.1's capability at 30% of the inference cost.

Tool CallingVisionLong Context

Price

$1.000 avg / 1M tokens

Context

262144

See gpt-4.1-mini pricing

gpt-5-mini

openai

OpenAI's most compact and cost-effective GPT-5 model, released August 7, 2025. Designed for high-volume, latency-sensitive applications where the full capabilities of larger models are not required, while maintaining strong performance on common tasks.

Tool CallingLong Context

Price

$1.125 avg / 1M tokens

Context

400K

See gpt-5-mini pricing

deepseek-reasoner

deepseek

DeepSeek-Reasoner is a 340B-parameter hybrid MoE model specialized in multi-step logical reasoning, causal inference, and problem decomposition. Features chain-of-thought verification, uncertainty quantification, and tool-integrated reasoning across mathematical, scientific, and real-world decision-making domains.

Tool CallingLong Context

Price

$1.165 avg / 1M tokens

Context

256K

See deepseek-reasoner pricing

moonshotai/Kimi-K2-Instruct

novita

Kimi-K2-Instruct is Moonshot AI's flagship 100-billion-parameter instruction-tuned model featuring a hybrid dense-MoE architecture with 200K context handling. Optimized for complex reasoning, multilingual dialogue, and long-form document understanding. Specializes in Chinese and English technical domains with enhanced safety alignment.

Tool CallingLong Context

Price

$1.435 avg / 1M tokens

Context

128K

See moonshotai/Kimi-K2-Instruct pricing

moonshotai/kimi-k2-0905

novita

Kimi-K2-0905 is a 72B-parameter multimodal language model optimized for long-context understanding and complex reasoning. Features enhanced multilingual capabilities, advanced tool usage, and strong performance in mathematical and logical reasoning tasks with extended context handling.

Tool CallingVisionLong Context

Price

$1.550 avg / 1M tokens

Context

128K

See moonshotai/kimi-k2-0905 pricing

deepseek-ai/DeepSeek-Prover-V2-671B

novita

DeepSeek-Prover-v2-671b is a 671B-parameter hybrid dense-MoE model specialized in mathematical theorem proving and formal reasoning, featuring 128 experts with 4-6 active per token. Optimized with FP8 quantization for efficient large-scale inference.

Price

$1.600 avg / 1M tokens

Context

32768

See deepseek-ai/DeepSeek-Prover-V2-671B pricing

Qwen/Qwen3-235B-A22B-Thinking-2507

novita

Qwen3-235B-A22B-Thinking-2507 is Alibaba Cloud's cognitive architecture model specializing in advanced reasoning, theory of mind simulation, and multi-agent collaboration. Built on the Qwen3-235B MoE foundation, it features specialized 'cognitive experts' and a neuro-symbolic execution engine for human-like problem-solving. Excels at strategic planning, counterfactual reasoning, and complex system modeling with 256K context retention.

Tool CallingLong Context

Price

$1.650 avg / 1M tokens

Context

256K

See Qwen/Qwen3-235B-A22B-Thinking-2507 pricing

moonshotai/Kimi-K2.5

novita

Kimi K2.5 is a large open-weight MoE model with 1.1 trillion parameters, known for strong vision capabilities via MoonViT-3D and agent swarm orchestration.

Tool CallingVisionLong Context

Price

$1.800 avg / 1M tokens

Context

256K

See moonshotai/Kimi-K2.5 pricing

moonshotai/Kimi-K2.6

novita

Kimi-K2.6 is an open-source native multimodal agentic model developed by Moonshot AI, supporting long-horizon coding capabilities and agentic task orchestration scaling to 300 sub-agents.

Tool CallingVisionLong Context

Price

$2.100 avg / 1M tokens

Context

256K

See moonshotai/Kimi-K2.6 pricing

zai-org/GLM-5

novita

GLM-5 is a 754B-parameter Mixture-of-Experts model from Z.ai. It features DeepSeek Sparse Attention (DSA) for efficient long-context handling, leading open-source models in long-horizon planning and agentic tasks.

Tool CallingLong Context

Price

$2.100 avg / 1M tokens

Context

200K

See zai-org/GLM-5 pricing

deepseek-ai/DeepSeek-V4-Pro

novita

DeepSeek-V4-Pro is a flagship Mixture-of-Experts (MoE) language model with 862B total parameters (49B activated) and a 1-million-token context window. It features a Hybrid Attention Architecture combining Compressed Sparse Attention and Heavily Compressed Attention for efficient long-context processing. With multiple distinct reasoning modes (such as Think Max for maximum reasoning effort), it is built for advanced mathematical reasoning, software engineering, tool use scenarios, and complex long-horizon agentic workflows.

Tool CallingLong Context

Price

$2.400 avg / 1M tokens

Context

1M

See deepseek-ai/DeepSeek-V4-Pro pricing

moonshotai/Kimi-K2.7-Code

novita

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Tool CallingVisionLong Context

Price

$2.475 avg / 1M tokens

Context

256K

See moonshotai/Kimi-K2.7-Code pricing

gpt-5.4-mini

openai

GPT-5.4 Mini is a compact, cost-efficient model designed for reliable performance across high-volume, everyday AI workloads. It offers fast, efficient reasoning, strong multimodal capabilities, and approaches the performance of the larger GPT-5.4 model on coding and agentic tasks while running more than 2x faster than its predecessor.

Tool CallingVisionLong Context

Price

$2.625 avg / 1M tokens

Context

400K

See gpt-5.4-mini pricing

zai-org/GLM-5.2

novita

GLM-5.2 is Z.ai's latest flagship open-weights MoE model engineered specifically to dominate long-horizon autonomous coding and engineering tasks. It features a solid 1-million-token context window, multiple thinking effort levels, and an improved IndexShare architecture that significantly boosts reasoning and agentic performance.

Tool CallingLong Context

Price

$2.900 avg / 1M tokens

Context

1M

See zai-org/GLM-5.2 pricing

claude-sonnet-5

anthropic

Claude Sonnet 5 is the most agentic Sonnet model yet, released by Anthropic on June 30, 2026. It offers a 1M token context window and 128k max output tokens. It is designed for complex, multi-step agentic tasks like coding, tool use, and autonomous planning, achieving performance close to the more expensive Opus 4.8 model. It features adaptive thinking enabled by default and a new tokenizer that produces approximately 30% more tokens for the same text compared to Sonnet 4.6.

Tool CallingVisionLong Context

Price

$3.000 avg / 1M tokens

Context

1M

See claude-sonnet-5 pricing

gpt-5.6-luna

openai

GPT-5.6 Luna is the fastest, most cost-efficient model in the GPT-5.6 family. It is designed for high-volume, latency-sensitive workloads such as classification, data extraction, request routing, and draft generation. It shares the same 1.05M context window as its siblings

Tool CallingVisionLong Context

Price

$3.500 avg / 1M tokens

Context

1.05M

See gpt-5.6-luna pricing

gpt-4.1

openai

GPT-4.1 is OpenAI's 1.6 trillion parameter multimodal reasoning system featuring a 256-expert MoE architecture with enhanced tool integration and world knowledge grounding. Optimized for enterprise applications with improved factual accuracy and reduced hallucination.

Tool CallingVisionLong Context

Price

$5.000 avg / 1M tokens

Context

524288

See gpt-4.1 pricing

gpt-5

openai

A large, multimodal language model with advanced reasoning capabilities and support for tools. It supports the new OpenAI Responses API with features like reasoning effort control.

Tool CallingVisionLong Context

Price

$5.625 avg / 1M tokens

Context

400K

See gpt-5 pricing

gpt-4o

openai

GPT-4o (Omnimodal) is OpenAI's 1.2 trillion parameter multimodal foundation model featuring unified input processing across text, vision, and audio. Optimized for real-time interaction with enhanced reasoning and cross-modal understanding capabilities.

Tool CallingVisionLong Context

Price

$6.250 avg / 1M tokens

Context

128K

See gpt-4o pricing

gpt-5.2

openai

The standard 'Thinking' variant of OpenAI's GPT-5.2 model, released Dec 11, 2025. Excels at complex reasoning, coding, and multi-step projects. Optimized for professional knowledge work and serves as the balanced option between speed and capability in the GPT-5.2 family.

Tool CallingVisionLong Context

Price

$7.875 avg / 1M tokens

Context

400K

See gpt-5.2 pricing

gpt-5.4

openai

GPT-5.4 is OpenAI's flagship frontier model built for sustained, multi-step reasoning with reliable follow-through. It excels at agentic workflows, research, document analysis, and powering complex internal tools with robust instruction following.

Tool CallingVisionLong Context

Price

$8.750 avg / 1M tokens

Context

1M

See gpt-5.4 pricing

gpt-5.6-terra

openai

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, the balanced workhorse tier of the GPT-5.6 family. Delivers highly reliable performance for everyday professional workflows, code generation, and general-purpose agentic tasks at a lower cost than Sol.

Tool CallingVisionLong Context

Price

$8.750 avg / 1M tokens

Context

1M

See gpt-5.6-terra pricing

claude-sonnet-4-6

anthropic

Claude Sonnet 4.6 is Anthropic's balanced production model, offering near-Opus level intelligence at a much lower cost. It features a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work, and design, along with an adaptive thinking capability.

Tool CallingVisionLong Context

Price

$9.000 avg / 1M tokens

Context

1M

See claude-sonnet-4-6 pricing

Kimi-K3

novita

Kimi’s most capable flagship model to date, with 2.8 trillion parameters. Built for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. The world’s first open-source model in the 3-trillion-parameter class.

Tool CallingVisionLong Context

Price

$9.000 avg / 1M tokens

Context

1M

See Kimi-K3 pricing

claude-opus-4-6

anthropic

Claude Opus 4.6 brings Anthropic's advanced reasoning capabilities to high-stakes workflows. With full adaptive thinking support and a 1-million-token context window, it excels at complex tasks across coding, cybersecurity, financial analysis, and large-scale enterprise agents.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See claude-opus-4-6 pricing

claude-opus-4-7

anthropic

Claude Opus 4.7 excels in autonomous multi-step engineering and long-horizon agentic workflows. It features an Effort Control layer for self-correction and advanced visual intelligence, making it an ideal choice for software engineering, legal, and financial analyses.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See claude-opus-4-7 pricing

claude-opus-4-8

anthropic

Claude Opus 4.8 is a workhorse flagship model designed for high-end reasoning and long-horizon agentic coding. It offers exceptional consistency and autonomy to sustain work on complex, long-running tasks.

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See claude-opus-4-8 pricing

claude-opus-5

anthropic

Claude Opus 5 is Anthropic's flagship model designed for complex agentic coding and enterprise work, delivering intelligence close to Claude Fable 5 at half the price. It features a 1M token context window, adaptive thinking, and state-of-the-art performance on coding and knowledge work evaluations. It is optimized for bounded, complex tasks

Tool CallingVisionLong Context

Price

$15.000 avg / 1M tokens

Context

1M

See claude-opus-5 pricing

gpt-5.5

openai

GPT-5.5 is a highly capable and efficient model featuring dynamic routing between 'Instant' and 'Thinking' modes. It provides improvements in conceptual clarity, scientific research ability, accuracy during knowledge work, and polished conversational tones.

Tool CallingVisionLong Context

Price

$17.500 avg / 1M tokens

Context

1M

See gpt-5.5 pricing

gpt-5.6-sol

openai

GPT-5.6 Sol is the flagship model in the GPT-5.6 family, designed for the most complex reasoning, agentic coding, and deep research tasks. It features a 1.05M token context window and excels in long-horizon, multi-step workflows. It supports advanced features like programmatic tool calling, multi-agent orchestration, and a 'max' reasoning effort

Tool CallingVisionLong Context

Price

$17.500 avg / 1M tokens

Context

1M

See gpt-5.6-sol pricing

claude-fable-5

anthropic

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins.

Tool CallingVisionLong Context

Price

$30.000 avg / 1M tokens

Context

1M

See claude-fable-5 pricing

gpt-5-pro

openai

OpenAI's flagship GPT-5 Pro model, released October 6, 2025. A highly capable multimodal model featuring advanced reasoning, native tool use, and strong performance on complex professional and coding tasks. It introduced a structured 'reasoning effort' parameter for controlling response depth.

Tool CallingVisionLong Context

Price

$67.500 avg / 1M tokens

Context

400K

See gpt-5-pro pricing

gpt-5.4-pro

openai

GPT-5.4 Pro offers deeper, higher-reliability reasoning for complex production scenarios. It is engineered for high-stakes agentic workflows, long-form analysis and synthesis, complex planning, and advanced reasoning tasks where accuracy is paramount.

Tool CallingVisionLong Context

Price

$105.000 avg / 1M tokens

Context

1M

See gpt-5.4-pro pricing

gpt-5.5-pro

openai

GPT-5.5 Pro is the highest-capability model in the 5.5 lineup, optimized for the hardest tasks and long-running workflows. It uses significant compute to 'think' harder before answering, making it ideal for deep research, highly complex coding, and extensive multi-step problem solving.

Tool CallingVisionLong Context

Price

$105.000 avg / 1M tokens

Context

1M

See gpt-5.5-pro pricing

claude-3-haiku-20240307

6ce62836a581464f8ea1e6ee1d397a99

Claude 3 Haiku is Anthropic's fastest and most compact model in the Claude 3 family, optimized for speed and cost-effectiveness while maintaining strong performance for everyday tasks and high-volume operations.

Tool CallingVisionLong Context

Price

See providers

Context

200K

See claude-3-haiku-20240307 pricing

claude-haiku-4-5-20251001

1be62feac9264117aee04eb7cdc868f2

Claude Haiku 4.5 is Anthropic's fastest and most efficient model in the Claude 4 family, optimized for speed and cost-effectiveness while maintaining high quality performance for everyday tasks.

Tool CallingVision

Price

See providers

Context

N/A

See claude-haiku-4-5-20251001 pricing

claude-opus-4-1-20250805

f0f29ca2cd6c408c817f31623e97e73f

Claude Opus 4.1 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning and analysis tasks. Part of the earlier Claude 4 generation alongside Claude Opus 4.

Tool CallingVisionLong Context

Price

See providers

Context

200K

See claude-opus-4-1-20250805 pricing

claude-opus-4-20250514

9d44f0f5167143d6a783d42ca0ba64e4

Claude Opus 4 is a highly capable model in the Claude 4 family, offering strong performance for complex reasoning, analysis, and challenging tasks requiring deep understanding.

Tool CallingVisionLong Context

Price

See providers

Context

200K

See claude-opus-4-20250514 pricing

claude-opus-4-5-20251101

14991ae9ac7d45298b4e7e05db98b4fc

Claude Opus 4.5 is Anthropic's most capable model in the Claude 4 family, offering superior performance for complex tasks requiring deep reasoning and analysis.

Tool CallingVision

Price

See providers

Context

N/A

See claude-opus-4-5-20251101 pricing

claude-sonnet-4-20250514

91cad413eb7148658649629510425980

Claude Sonnet 4 is a balanced model in the Claude 4 family, offering a strong combination of intelligence, speed, and efficiency for a wide range of tasks including reasoning, coding, and analysis.

Tool CallingVisionLong Context

Price

See providers

Context

200K

See claude-sonnet-4-20250514 pricing

claude-sonnet-4-5-20250929

a40f171ceba14d75b2bf1749904d0965

Claude Sonnet 4.5 is Anthropic's smartest model and the flagship of the Claude 4 family, offering the best balance of intelligence, speed, and efficiency for everyday use. It excels at complex reasoning, coding, and analysis tasks.

Tool CallingVision

Price

See providers

Context

N/A

See claude-sonnet-4-5-20250929 pricing

deepseek/deepseek-r1-distill-llama-8b

24c0b83803ff4ac8b32693aee54098e8

DeepSeek-R1-Distill-Llama-8B combines DeepSeek-R1 knowledge distillation with Llama architecture. Available in FP8 (H100+ only) and FP16 quantization, delivering efficient performance with 32K context.

Tool Calling

sarvamai/sarvam-105b

sarvam

Sarvam-105B is an advanced Mixture-of-Experts (MoE) model with 10.3B active parameters, designed for superior performance across a wide range of complex tasks. It is highly optimized for complex reasoning, with particular strength in agentic tasks, mathematics, and coding.

Tool CallingLong Context

Price

See providers

Context

128K

See sarvamai/sarvam-105b pricing

How to choose

Conversational applications
AI agents
Code generation
RAG and document analysis

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How are text model prices compared?

When input and output token prices are available, OneInfer uses their average as a simple comparison score. The full input and output rates remain visible on each model page.

Can I switch text models without changing my integration?

Yes. Models are accessed through the unified OneInfer API; change the model identifier while keeping the same authentication and request pattern.