Embedding and retrieval APIs

Compare Embedding Models and API Pricing

Find embedding models for semantic search, retrieval-augmented generation, recommendations, clustering, and document processing through a consistent API.

Current catalog

1

embedding models with verified catalog metadata

Pricing content refreshed July 25, 2026

Available Embedding Models

Comparable models are ordered by their current normalized price.

1 models

sarvamai/sarvam-30b

sarvam

An advanced Mixture-of-Experts (MoE) model with 2.4B non-embedding active parameters, designed for practical deployment. It combines strong reasoning, reliable coding ability, and best-in-class conversational quality across 22 Indian languages, and supports tool calling in multilingual voice call scenarios.

Tool CallingEmbeddings

Price

See providers

Context

64K

See sarvamai/sarvam-30b pricing

How to choose

RAG
Semantic search
Recommendations
Document clustering

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How are embedding model prices compared?

Embedding models are compared using their positive input-token price per one million tokens when that pricing is available.

Can I use these models with my existing vector database?

Yes. Store the returned vectors in your preferred vector database and use them for retrieval or similarity search.