ModelsAgree
← All leaderboards
🎙

Best speech-to-text API

4 models · updated 2026-07-15

The verdict

Deepgram leads — All 4 models rank Deepgram the top pick.

As of 2026-07-15, ChatGPT, Claude, Gemini and Grok collectively rank Deepgram #1 for speech-to-text api on ModelsAgree — a unanimous pick. The models' case: Best overall balance of streaming accuracy, very low latency, competitive cost, diarization, multilingual support, and production-ready controls for noisy calls and. The models' main caveat: Accuracy varies by language and accent, so multilingual teams must benchmark their actual traffic. The strongest alternative is AssemblyAI — Universal-2 sits at or near the top of independent accuracy benchmarks with the strongest bundled audio intelligence — speaker diarization, PII. Source: https://modelsagree.com/best/best-speech-to-text-api (modelsagree.com, CC BY 4.0).

Grade any brand's AI visibility →See how ChatGPT, Claude, Gemini & Grok rate any product, or your own.

Combined ranking

  1. 1
    GPT #1Claude #1Gemini #1Grok #1

    Best overall balance of streaming accuracy, very low latency, competitive cost, diarization, multilingual support, and production-ready controls for noisy calls and domain vocabulary

    + model takes & fixes

    GPT Best overall balance of streaming accuracy, very low latency, competitive cost, diarization, multilingual support, and production-ready controls for noisy calls and domain vocabulary

    Claude Nova-3 hits the best speed/accuracy/price balance in the category — sub-300ms streaming latency, ~$0.0043/min batch pricing, keyterm prompting, and a self-hosted/container option; assumes the typical practitioner is a developer shipping real-time or high-volume transcription where cost per minute and latency dominate. Near-tie with AssemblyAI for batch-centric workloads.

    Gemini Industry-leading speed and sub-300ms latency, making it the definitive choice for real-time conversational voice agents, combined with a highly cost-effective pricing model and developer-friendly streaming SDKs. (Assuming real-time latency and raw cost-efficiency are prioritized over advanced post-transcript LLM analysis.)

    Grok Lowest real-time latency with fast EOT detection, strong accuracy on noisy/conversational audio, competitive pricing (~$0.26/hr batch), excellent for voice agents and high-volume streaming with good multilingual support

    Where it falls short

    per GPT Accuracy varies by language and accent, so multilingual teams must benchmark their actual traffic

    per Claude raw accuracy on heavily accented, noisy, or long-tail multilingual audio trails Whisper-large and Scribe — not the pick if maximum WER on hard audio matters more than latency and cost.

    per Gemini Lacks native advanced audio intelligence LLM features out of the box, requiring developers to orchestrate external LLM pipelines for deep analysis, summarization, or semantic extraction.

    per Grok Broaden native multilingual code-switching and add deeper built-in speech intelligence (sentiment/entities) without extra cost

  2. 2
    GPT #3Claude #2Gemini #2Grok #2

    Universal-2 sits at or near the top of independent accuracy benchmarks with the strongest bundled audio intelligence — speaker diarization, PII redaction, sentiment, chapters — plus clean docs and SDKs, so one API covers what elsewhere takes three services.

    + model takes & fixes

    Claude Universal-2 sits at or near the top of independent accuracy benchmarks with the strongest bundled audio intelligence — speaker diarization, PII redaction, sentiment, chapters — plus clean docs and SDKs, so one API covers what elsewhere takes three services.

    Gemini Unmatched out-of-the-box Audio Intelligence suite (including LeMUR for LLM-powered summarization, speaker diarization, and PII redaction natively integrated) and superior accuracy in specialized technical/medical domains. Near-tie with OpenAI Whisper for batch processing, but wins on feature completeness.

    Grok Top-tier accuracy especially on challenging real-world audio with fewer hallucinations, rich built-in audio intelligence (diarization, topics, PII), strong developer tools for voice agents and analytics

    GPT Excellent developer experience plus strong transcription, diarization, formatting, language detection, and speech-intelligence features that reduce downstream engineering

    Where it falls short

    per GPT Not the best value when only basic high-volume transcription is needed

    per Claude cloud-only with no self-hosted deployment, and real-time streaming is weaker and more English-centric than Deepgram's — wrong choice for on-prem/compliance-bound or latency-critical voice agents.

    per Gemini Premium cost structure is significantly more expensive at scale compared to raw transcript competitors, and latency is too high for real-time voice agents.

    per Grok Further reduce streaming latency to sub-200ms consistently and expand real-time language/code-switching coverage

  3. 3
    GPT #4Claude #3Gemini #3Grok

    the open-source default — free weights, 99 languages, and a huge ecosystem (faster-whisper, whisper.cpp, WhisperX) that makes self-hosting cheap at scale and keeps audio in-house; still the best value when you have GPUs and engineering time.

    + model takes & fixes

    Claude the open-source default — free weights, 99 languages, and a huge ecosystem (faster-whisper, whisper.cpp, WhisperX) that makes self-hosting cheap at scale and keeps audio in-house; still the best value when you have GPUs and engineering time.

    Gemini The global gold standard for out-of-the-box multilingual accuracy and translation capabilities (direct-to-English) supported by a massive developer ecosystem, allowing teams to choose between the managed API or self-hosted open-source model.

    GPT Particularly strong on difficult accents, noisy audio, and terminology when supplied with context; simple API and attractive accuracy-per-dollar for file transcription

    Where it falls short

    per GPT Fewer mature speech-specific controls and deployment options than established STT platforms

    per Claude no native streaming or diarization — you stitch those on yourself (WhisperX/pyannote), run your own inference ops, and manage its known hallucinations on silence and music.

    per Gemini High latency and lack of native support for essential transcription features like speaker diarization and PII redaction, which must be built manually.

  4. 4
    GPT Claude #4Gemini Grok #3

    Excellent multilingual accuracy across 90+ languages with low latency realtime, seamless integration for full voice pipelines (with their TTS), strong on code-switching and production audio

    + model takes & fixes

    Grok Excellent multilingual accuracy across 90+ languages with low latency realtime, seamless integration for full voice pipelines (with their TTS), strong on code-switching and production audio

    Claude launched 2025 straight to the top of multilingual WER benchmarks (FLEURS, Common Voice), with word-level timestamps, diarization, and audio-event tagging across ~99 languages — the accuracy leader for batch transcription of hard, multilingual audio.

    Where it falls short

    per Claude batch-first product — its real-time offering is newer and less proven, and per-minute pricing runs higher than Deepgram at volume, so it's not the pick for cost-sensitive streaming.

    per Grok Improve cost-efficiency for high-volume usage and expand on-prem/self-hosted deployment options

  5. 5
    GPT #5Claude #5Gemini #4Grok

    The undisputed leader for enterprise deployments that require strict data sovereignty, offering fully air-gapped on-premises or private cloud deployments combined with robust support for domain-tuned models.

    + model takes & fixes

    Gemini The undisputed leader for enterprise deployments that require strict data sovereignty, offering fully air-gapped on-premises or private cloud deployments combined with robust support for domain-tuned models.

    GPT Strong real-world multilingual and accented-speech recognition, capable streaming, diarization, and flexible cloud or self-hosted enterprise deployment

    Claude the accent- and dialect-robustness leader — consistently strongest on non-native and regional English plus solid 50-language coverage, with genuine deployment flexibility (SaaS, container, on-prem) that enterprises with data-residency constraints need.

    Where it falls short

    per GPT Pricing and onboarding are less transparent and self-serve-friendly than the leaders

    per Claude costs more and the developer experience is less polished than the dev-first APIs above — overkill for a typical startup that just needs good English transcription fast.

    per Gemini Extremely high cost of entry and complex enterprise sales cycles, making it completely over-engineered for small projects or early-stage startups.

  6. 6
    GPT #2Claude Gemini Grok

    Near-tied for first on recognition quality, with broad language coverage, strong streaming and batch modes, speaker features, and mature enterprise infrastructure

    + model takes & fixes

    GPT Near-tied for first on recognition quality, with broad language coverage, strong streaming and batch modes, speaker features, and mature enterprise infrastructure

    Where it falls short

    per GPT Google Cloud configuration, quotas, regions, and pricing are more cumbersome than specialist APIs

  7. 7
    GPT Claude Gemini #5Grok #5

    Exceptional at handling complex, noisy real-world audio and multi-language code-switching (mixed languages mid-sentence) with a pricing model that bundles features like speaker diarization and language detection at no extra cost.

    + model takes & fixes

    Gemini Exceptional at handling complex, noisy real-world audio and multi-language code-switching (mixed languages mid-sentence) with a pricing model that bundles features like speaker diarization and language detection at no extra cost.

    Grok Leading accuracy on noisy real-world business/conversational audio in core languages, competitive pricing with generous free tier, strong multilingual and code-switching capabilities

    Where it falls short

    per Gemini Lacks the extensive community support and developer ecosystem maturity of OpenAI or Deepgram, with no robust on-premises deployment tier.

    per Grok Enhance developer ecosystem integrations and advanced speech understanding features like native sentiment/topic detection

  8. 8
    GPT Claude Gemini Grok #4

    Outstanding overall accuracy for batch processing, robust handling of accents/noise/technical vocab, broad language support and ecosystem integration

    + model takes & fixes

    Grok Outstanding overall accuracy for batch processing, robust handling of accents/noise/technical vocab, broad language support and ecosystem integration

    Where it falls short

    per Grok Significantly lower realtime latency and pricing for streaming/high-volume production use cases

By use case

How this board's leaders rank when the same four models are asked a more specific question.

ProductThis boardreal-timecall centerscheapmedical transcription
Deepgram#1#1#1#4#2
AssemblyAI#2#2#2#3#4
OpenAI#3
ElevenLabs#4
Speechmatics#5#3#4#8#10
Google Cloud Speech-to-Text#6#6#6#10
Gladia#7#5#5
OpenAI Whisper#8#9

Rank history

1234567891006-2907-0807-1007-1307-15DeepgramAssemblyAIOpenAIElevenLabsSpeechmaticsGoogle Cloud Speech-to-TextGladiaOpenAI Whisper
Deepgram#1AssemblyAI#2OpenAI#3ElevenLabs#4Speechmatics#7Google Cloud Speech-to-Text#5Gladia#8OpenAI Whisper#6

Just missed the top 5

GPT Azure AI Speechbroad, mature, and customizable, but operational complexity and inconsistent language-by-language quality hold it back · Amazon Transcribereliable AWS-native choice, but generally less compelling on accuracy, developer experience, and value than the top five

Claude Google Cloud Speech-to-TextChirp 2 is competent and convenient inside GCP, but pricing, DX, and accuracy don't beat the specialists at anything

Gemini Groq Whisper APIoffers incredibly fast and cheap inference for standard Whisper models, but missed the top 5 due to a total lack of auxiliary features like diarization or formatting control · whisper.cpphighly optimized C/C++ port of Whisper that is perfect for local/on-device transcription, but requires manual infrastructure management and is not a managed API out of the box

Grok Speechmaticsstrong accuracy and flexible on-prem deployment but lags in realtime latency and cost for agents · Google Cloud Speech-to-Textbroad ecosystem and languages but lower accuracy/latency in independent benchmarks

By model

ChatGPT

  1. 1.Deepgram
  2. 2.Google Cloud Speech-to-Text
  3. 3.AssemblyAI
  4. 4.OpenAI
  5. 5.Speechmatics

Claude

  1. 1.Deepgram
  2. 2.AssemblyAI
  3. 3.OpenAI
  4. 4.ElevenLabs
  5. 5.Speechmatics

Gemini

  1. 1.Deepgram
  2. 2.AssemblyAI
  3. 3.OpenAI
  4. 4.Speechmatics
  5. 5.Gladia

Grok

  1. 1.Deepgram
  2. 2.AssemblyAI
  3. 3.ElevenLabs
  4. 4.OpenAI Whisper
  5. 5.Gladia

Common questions

What is the best speech-to-text api according to AI models?

Deepgram leads. All 4 models rank Deepgram the top pick. The current top 3: Deepgram, AssemblyAI, OpenAI. Ranked by asking ChatGPT, Claude, Gemini, Grok the same buying question and merging their top-5 picks, updated 2026-07-15. Source: modelsagree.com.

Which speech-to-text api did each AI model pick first?

ChatGPT: Deepgram. Claude: Deepgram. Gemini: Deepgram. Grok: Deepgram.

What changed in the latest speech-to-text api ranking?

In the latest poll (2026-07-15): ElevenLabs climbed 1 spot, Speechmatics climbed 4 spots, Gladia climbed 1 spot; OpenAI Whisper dropped 1 spot. The models are re-polled on demand, so this ranking moves.

How is this speech-to-text api ranking made?

ChatGPT, Claude, Gemini, Grok are each asked the same buying question in a fresh session with no system steering. Their top-5 answers are merged (rank 1 = 5 pts … rank 5 = 1 pt) into the consensus ranking, re-polled on demand and tracked over time.

More on how polling works: full methodology →

Cite this ranking

ModelsAgree, “Best speech-to-text API” — merged ranking from ChatGPT, Claude, Gemini & Grok, polled 2026-07-15. https://modelsagree.com/best/best-speech-to-text-api (CC BY 4.0)

Tracked by ModelsAgree · rank 1 = 5 pts … rank 5 = 1 pt · re-polled on demand