Transcription APIs

Compare Speech-to-Text APIs and Transcription Pricing

Find automatic speech recognition models for meetings, podcasts, voice notes, call analytics, and real-time transcription. Per-minute prices are ranked separately from token-based audio prices.

Current catalog

15

speech-to-text models with verified catalog metadata

Pricing content refreshed September 8, 2026

Lowest comparable price

Google: Gemini 2.5 Flash Lite

Currently the lowest-priced option in the average of available input and output prices per one million tokens group. Different billing units are never mixed.

$0.250 avg / 1M tokens

View pricing and providers

Available Speech-to-Text Models

Comparable models are ordered by their current normalized price.

15 models

Google: Gemini 2.5 Flash Lite

openrouter

Lowest price

Gemini 2.5 Flash-Lite is Google's lightweight and cost-efficient reasoning model in the Gemini 2.5 family, optimized for high-throughput, low-latency workloads. It supports multimodal inputs including text, images, audio, video, and PDFs, with optional thinking for applications that need a balance between speed, cost, and reasoning quality.

Tool CallingVisionAudio InputSpeech To Text

Price

$0.250 avg / 1M tokens

Context

1M

See Google: Gemini 2.5 Flash Lite pricing

Google: Gemini 3.1 Flash-Lite

openrouter

Gemini 3.1 Flash-Lite is Google's low-latency, cost-effective multimodal model optimized for high-frequency and high-volume workloads. It supports text, image, video, audio, and PDF inputs and is designed for lightweight reasoning, data extraction, agentic workflows, tool use, and applications where latency and API cost are primary considerations.

Tool CallingVisionAudio InputSpeech To Text

Price

$0.875 avg / 1M tokens

Context

1M

See Google: Gemini 3.1 Flash-Lite pricing

Google: Gemini 2.5 Flash

openrouter

Gemini 2.5 Flash is Google's high-performance multimodal workhorse model designed for reasoning, coding, mathematics, science, document understanding, and agentic applications while maintaining Flash-tier speed and cost efficiency.

Tool CallingVisionAudio InputSpeech To Text

Price

$1.400 avg / 1M tokens

Context

1M

See Google: Gemini 2.5 Flash pricing

Google: Gemini 3.5 Flash Lite

openrouter

Gemini 3.5 Flash Lite is Google's high-efficiency multimodal model with upgraded agentic capabilities, optimized for low-cost, high-volume workloads and focused subagents operating inside complex multi-agent systems.

Tool CallingVisionAudio InputSpeech To Text

Price

$1.400 avg / 1M tokens

Context

1M

See Google: Gemini 3.5 Flash Lite pricing

Google: Gemini 3.6 Flash

openrouter

Gemini 3.6 Flash is a frontier-class Google multimodal Flash model designed for fast reasoning, coding, multimodal understanding, tool use, and scalable agentic applications while retaining Flash-tier latency and efficiency.

Tool CallingVisionAudio InputSpeech To Text

Price

$2.250 avg / 1M tokens

Context

1M

See Google: Gemini 3.6 Flash pricing

Google: Gemini 3.7 Flash

openrouter

Gemini 3.7 Flash is a frontier-class Google Flash model focused on high-performance reasoning, coding, multimodal understanding, tool use, and agentic execution with the latency and cost characteristics of the Flash family.

Tool CallingVisionAudio InputSpeech To Text

Price

$2.250 avg / 1M tokens

Context

1M

See Google: Gemini 3.7 Flash pricing

Google: Gemini 3.8 Flash

openrouter

Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, complex enterprise workflows, advanced reasoning, and multimodal understanding. Building on Gemini 3.7 Flash, it delivers substantial improvements across coding, agentic tasks, and critical multi-step reasoning while retaining Flash-level speed and cost efficiency. It supports a 1M-token context window, configurable thinking levels, function calling, computer use, code execution, search grounding, structured outputs, file search, and URL context.

Tool CallingVisionAudio InputSpeech To Text

Price

$2.250 avg / 1M tokens

Context

1M

See Google: Gemini 3.8 Flash pricing

Meta: Muse Spark 1.1

meta

Muse Spark 1.1 is Meta's flagship proprietary multimodal frontier model developed by Meta Superintelligence Labs (MSL). Designed for complex reasoning and agentic tasks, it can actively manage a 1-million-token context window, orchestrate multi-agent systems, and natively process text, image, video, and audio inputs. It excels at long-horizon software development, tool use, and computer-use workflows.

Tool CallingVisionAudio InputSpeech To Text

Price

$2.750 avg / 1M tokens

Context

1M

See Meta: Muse Spark 1.1 pricing

Meta: Muse Spark 1.2

meta

Muse Spark 1.2 is a proprietary coding-focused frontier model developed by Meta Superintelligence Labs. Co-trained specifically with the Muse Code terminal agent, it features significant improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. It leverages planning, goal conditioning, and context compaction to sustain progress on long-horizon tasks like whole-repository generation and GPU kernel optimization.

Tool CallingVisionAudio InputSpeech To Text

Price

$2.750 avg / 1M tokens

Context

1M

See Meta: Muse Spark 1.2 pricing

Google: Gemini 3.5 Flash

openrouter

Gemini 3.5 Flash is Google's high-efficiency multimodal model offering near-Pro-level coding and reasoning at Flash-tier speed and cost. It is optimized for coding, parallel agentic execution, multimodal understanding, and scalable production workloads.

Tool CallingVisionAudio InputSpeech To Text

Price

$5.250 avg / 1M tokens

Context

1M

See Google: Gemini 3.5 Flash pricing

Google: Gemini 2.5 Pro

openrouter

Gemini 2.5 Pro is Google's advanced Gemini 2.5 reasoning model for complex problem solving, coding, mathematics, scientific analysis, long-context understanding, multimodal reasoning, and sophisticated agentic workflows.

Tool CallingVisionAudio InputSpeech To Text

Price

$5.625 avg / 1M tokens

Context

1M

See Google: Gemini 2.5 Pro pricing

Google: Gemini 3.1 Pro

openrouter

Gemini 3.1 Pro is Google's advanced intelligence model for complex reasoning, problem solving, software engineering, multimodal analysis, long-context tasks, and powerful agentic and vibe-coding workflows.

Tool CallingVisionAudio InputSpeech To Text

Price

$7.000 avg / 1M tokens

Context

1M

See Google: Gemini 3.1 Pro pricing

Fishaudio: S2 Pro

openrouter

Fish Audio S2 Pro is an advanced 4B parameter open-weight text-to-speech model built on a dual-autoregressive architecture. It supports approximately 80 languages and introduces fine-grained, word-level inline tag control, allowing users to embed open-ended natural language instructions inside square brackets directly in the script (e.g., [whispering], [pitch up]). It outputs 44.1 kHz audio and natively supports zero-shot voice cloning from a reference clip.

Audio OutputSpeech To TextText To Speech

Price

See providers

Context

4K

See Fishaudio: S2 Pro pricing

Sarvam AI: Saaras V3

sarvam

Saaras v3 is a state-of-the-art speech recognition and translation model. It natively supports streaming for real-time applications and features specialized modes for transcription, direct-to-English translation, and transliteration. It is specifically tuned for noisy environments and complex code-mixed (e.g., Hindi-English) speech.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See Sarvam AI: Saaras V3 pricing

Sarvam AI: Saarika V2.5

sarvam

Saarika v2.5 is Sarvam AI's legacy speech recognition model designed for Indian languages and accents. It transcribes audio in the same language spoken, excelling in multi-speaker conversations, telephony audio (8kHz), and code-mixed speech. Supports 11 languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, English) with automatic language detection and speaker diarization. Achieves 4.96% CER and 18.32% WER on VISTAAR benchmark.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See Sarvam AI: Saarika V2.5 pricing

How to choose

Meeting transcription
Podcast transcription
Call analytics
Voice interfaces

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How is the cheapest speech-to-text API selected?

OneInfer compares current positive per-audio-minute prices when providers expose that unit. Models using another billing unit are shown but not included in that ranking.

Do all transcription models provide timestamps?

No. Timestamp support is a model capability and should be confirmed on the individual model page.