Transcription APIs

Compare Speech-to-Text APIs and Transcription Pricing

Find automatic speech recognition models for meetings, podcasts, voice notes, call analytics, and real-time transcription. Per-minute prices are ranked separately from token-based audio prices.

Current catalog

4

speech-to-text models with verified catalog metadata

Pricing content refreshed July 25, 2026

Lowest comparable price

gpt-5.1

Currently the lowest-priced option in the average of available input and output prices per one million tokens group. Different billing units are never mixed.

$5.625 avg / 1M tokens

View pricing and providers

Available Speech-to-Text Models

Comparable models are ordered by their current normalized price.

4 models

gpt-5.1

openai

Lowest price

OpenAI's GPT-5.1 model, released November 13, 2025. Features a strong multimodal and reasoning capabilities, and is optimized for complex agentic workflows. It introduced significant improvements in coding, reasoning, and tool use over previous generations.

Tool CallingVisionAudio InputSpeech To Text

Price

$5.625 avg / 1M tokens

Context

400K

See gpt-5.1 pricing

gpt-5.2-pro

e4863abbf36c4fc8b10e095cc1aa9b6d

OpenAI's flagship GPT-5.2 Pro model, Designed for professional knowledge work and agentic workflows, exclusive 'xhigh' reasoning effort, and top-tier performance on complex reasoning, coding, and scientific benchmarks.

Tool CallingVisionAudio InputSpeech To Text

Price

See providers

Context

400K

See gpt-5.2-pro pricing

saaras:v3

sarvam

Saaras v3 is a state-of-the-art speech recognition and translation model. It natively supports streaming for real-time applications and features specialized modes for transcription, direct-to-English translation, and transliteration. It is specifically tuned for noisy environments and complex code-mixed (e.g., Hindi-English) speech.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See saaras:v3 pricing

saarika:v2.5

sarvam

Saarika v2.5 is Sarvam AI's legacy speech recognition model designed for Indian languages and accents. It transcribes audio in the same language spoken, excelling in multi-speaker conversations, telephony audio (8kHz), and code-mixed speech. Supports 11 languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, English) with automatic language detection and speaker diarization. Achieves 4.96% CER and 18.32% WER on VISTAAR benchmark.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See saarika:v2.5 pricing

How to choose

Meeting transcription
Podcast transcription
Call analytics
Voice interfaces

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How is the cheapest speech-to-text API selected?

OneInfer compares current positive per-audio-minute prices when providers expose that unit. Models using another billing unit are shown but not included in that ranking.

Do all transcription models provide timestamps?

No. Timestamp support is a model capability and should be confirmed on the individual model page.