{"slug":"deepgram","name":"Deepgram","domain":"deepgram.com","verdict":"As of 2026-07-13, ChatGPT, Claude, Gemini, Grok collectively rank Deepgram first for ai transcription api (one of 8 leaderboards it appears on). Source: https://modelsagree.com/product/deepgram (modelsagree.com, CC BY 4.0).","best_rank":1,"categories":8,"brief":{"category":"best-ai-transcription-api","title":"Best AI transcription API","rank":1,"of":8,"top":null,"day":"2026-07-16","why":[{"t":"excellent noisy-audio accuracy","m":["ChatGPT","Grok"],"q":"excellent noisy-audio accuracy"},{"t":"ultra-low latency for real-time streaming","m":["ChatGPT","Claude","Gemini","Grok"],"q":"ultra-low latency for real-time streaming"},{"t":"turn-taking detection for voice agents","m":["ChatGPT","Gemini"],"q":"built-in turn-taking detection via the Flux model"},{"t":"competitive pricing and clean SDKs","m":["Claude","Grok"],"q":"competitive pricing (~$0.004-0.014/min)"}],"gap":[],"fix":[{"t":"broaden native high-accuracy languages","m":["ChatGPT","Claude","Gemini","Grok"],"q":"Broaden native high-accuracy languages beyond English-dominant"},{"t":"audio-intelligence layer is thinner","m":["Claude"],"q":"Beyond core STT its audio-intelligence layer is thinner than AssemblyAI's"},{"t":"high-volume pricing scales steeply","m":["Gemini"],"q":"High-volume pricing scales steeply"}]},"entries":[{"slug":"best-ai-transcription-api","title":"Best AI transcription API","rank":1,"of":8,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Nova-3 combines excellent noisy-audio accuracy, fast batch and streaming transcription, precise timestamps, diarization, redaction, and effective keyterm prompting; Flux adds model-native turn detection for voice agents. Best overall assuming live and recorded workloads both matter.","reasons":[{"model":"ChatGPT","reason":"Nova-3 combines excellent noisy-audio accuracy, fast batch and streaming transcription, precise timestamps, diarization, redaction, and effective keyterm prompting; Flux adds model-native turn detection for voice agents. Best overall assuming live and recorded workloads both matter."},{"model":"Claude","reason":"The purpose-built STT specialist that best fits the typical developer: Nova-3 accuracy at aggressive per-minute pricing (~$0.004–0.005/min), sub-300ms streaming that has become the default for voice-agent stacks, keyterm prompting, self-hosted/VPC deployment, and clean SDKs and docs; assumption: the typical practitioner is a product developer who weighs latency, cost, and integration speed over exotic features."},{"model":"Gemini","reason":"Best-in-class real-time streaming with sub-300ms latency and built-in turn-taking detection via the Flux model, making it the top choice for voice agents."},{"model":"Grok","reason":"Leads benchmarks in accuracy (lowest WER on noisy/conversational audio), ultra-low latency for real-time streaming, competitive pricing (~$0.004-0.014/min), strong diarization and multilingual support for production voice agents."}],"fixes":[{"model":"ChatGPT","fix":"Nova-3’s strongest multilingual mode covers fewer languages than the broadest rivals, and advanced self-hosting is enterprise-oriented."},{"model":"Claude","fix":"Beyond core STT its audio-intelligence layer is thinner than AssemblyAI's, and accuracy on heavily accented or noisy multilingual audio can trail specialists like Speechmatics."},{"model":"Gemini","fix":"High-volume pricing scales steeply and its batch accuracy on niche multilingual dialects lags behind dedicated multilingual engines."},{"model":"Grok","fix":"Broaden native high-accuracy languages beyond English-dominant to match top multilingual rivals without fallback degradation"}],"updated":"2026-07-13","rank_history":{"days":["2026-07-11","2026-07-12","2026-07-13"],"ranks":[1,1,1]},"reasoning_shift":[{"model":"Gemini","from":"2026-07-12","to":"2026-07-13","added":[{"t":"built-in turn-taking detection","q":"built-in turn-taking detection via the Flux model"},{"t":"High-volume pricing scales steeply","q":"High-volume pricing scales steeply"},{"t":"batch accuracy on niche multilingual dialects","q":"its batch accuracy on niche multilingual dialects lags behind dedicated multilingual engines"}],"dropped":[{"t":"high accuracy-to-cost ratio","q":"high accuracy-to-cost ratio"},{"t":"robust developer APIs","q":"robust developer APIs with extensive SDK support"},{"t":"structured data extraction","q":"Needs better out-of-the-box structured data extraction and LLM-driven audio intelligence tools comparable to AssemblyAI."}]},{"model":"Claude","from":"2026-07-12","to":"2026-07-13","added":[{"t":"default for voice-agent stacks","q":"sub-300ms streaming that has become the default for voice-agent stacks"},{"t":"keyterm prompting","q":"keyterm prompting"},{"t":"audio-intelligence layer is thinner","q":"its audio-intelligence layer is thinner than AssemblyAI's"}],"dropped":[{"t":"self-serve onboarding","q":"self-serve onboarding"}]}],"api":"https://modelsagree.com/api/v1/best/best-ai-transcription-api.json"},{"slug":"best-realtime-speech-to-text-api","title":"Best real-time speech-to-text API","rank":1,"of":9,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Best overall balance of low latency, robust endpointing, noisy-audio and telephony accuracy, multilingual/code-switching support, mature WebSocket tooling, and competitive usage pricing; the safest default for production voice agents and live applications.","reasons":[{"model":"ChatGPT","reason":"Best overall balance of low latency, robust endpointing, noisy-audio and telephony accuracy, multilingual/code-switching support, mature WebSocket tooling, and competitive usage pricing; the safest default for production voice agents and live applications."},{"model":"Claude","reason":"Best overall latency/accuracy/price combination for live transcription — sub-300ms streaming over WebSocket, strong accuracy on real-world noisy audio, keyterm prompting for domain vocabulary, and pricing (~$0.0077/min streaming) well under the big-cloud providers; self-hosted/VPC deployment option covers regulated shops. The default pick for voice agents and live captioning unless you have a niche constraint."},{"model":"Gemini","reason":"Leading real-time speed with sub-300ms latency, highly cost-effective pricing, and advanced streaming-native features like native Voice Activity Detection (VAD) and integrated end-of-turn detection that optimize conversational agent responsiveness."},{"model":"Grok","reason":"Consistently tops or near-tops independent and company benchmarks for streaming accuracy (low WER on noisy/conversational audio), sub-300ms latency with strong turn detection/EOT in Flux, developer-friendly API, domain adaptation, and value for production voice agents/live apps; real-world production reliability and speed edge it for typical builders."}],"fixes":[{"model":"ChatGPT","fix":"Accuracy varies enough by accent and domain that specialized vocabulary or less-common languages require benchmarking and configuration."},{"model":"Claude","fix":"Language coverage and accented/low-resource-language accuracy trail Speechmatics and the hyperscalers — not the choice for heavily multilingual products."},{"model":"Gemini","fix":"Lacks a deeply integrated post-transcript analysis pipeline, requiring developers to chain external LLMs for structured data extraction."},{"model":"Grok","fix":"Not the absolute cheapest at high volume or broadest native languages without customization; streaming can have slight incremental vs. batch accuracy tradeoffs."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-realtime-speech-to-text-api.json"},{"slug":"best-speech-to-text-api","title":"Best speech-to-text API","rank":1,"of":8,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Best overall balance of streaming accuracy, very low latency, competitive cost, diarization, multilingual support, and production-ready controls for noisy calls and domain vocabulary","reasons":[{"model":"ChatGPT","reason":"Best overall balance of streaming accuracy, very low latency, competitive cost, diarization, multilingual support, and production-ready controls for noisy calls and domain vocabulary"},{"model":"Claude","reason":"Nova-3 hits the best speed/accuracy/price balance in the category — sub-300ms streaming latency, ~$0.0043/min batch pricing, keyterm prompting, and a self-hosted/container option; assumes the typical practitioner is a developer shipping real-time or high-volume transcription where cost per minute and latency dominate. Near-tie with AssemblyAI for batch-centric workloads."},{"model":"Gemini","reason":"Industry-leading speed and sub-300ms latency, making it the definitive choice for real-time conversational voice agents, combined with a highly cost-effective pricing model and developer-friendly streaming SDKs. (Assuming real-time latency and raw cost-efficiency are prioritized over advanced post-transcript LLM analysis.)"},{"model":"Grok","reason":"Lowest real-time latency with fast EOT detection, strong accuracy on noisy/conversational audio, competitive pricing (~$0.26/hr batch), excellent for voice agents and high-volume streaming with good multilingual support"}],"fixes":[{"model":"ChatGPT","fix":"Accuracy varies by language and accent, so multilingual teams must benchmark their actual traffic"},{"model":"Claude","fix":"raw accuracy on heavily accented, noisy, or long-tail multilingual audio trails Whisper-large and Scribe — not the pick if maximum WER on hard audio matters more than latency and cost."},{"model":"Gemini","fix":"Lacks native advanced audio intelligence LLM features out of the box, requiring developers to orchestrate external LLM pipelines for deep analysis, summarization, or semantic extraction."},{"model":"Grok","fix":"Broaden native multilingual code-switching and add deeper built-in speech intelligence (sentiment/entities) without extra cost"}],"updated":"2026-07-15","rank_history":{"days":["2026-06-29","2026-06-30","2026-07-08","2026-07-09","2026-07-10","2026-07-12","2026-07-13","2026-07-14","2026-07-15"],"ranks":[1,1,1,1,2,1,2,1,1]},"reasoning_shift":[{"model":"Gemini","from":"2026-07-14","to":"2026-07-15","added":[{"t":"Cost-effective pricing","q":"a highly cost-effective pricing model"},{"t":"Advanced audio intelligence missing","q":"Lacks native advanced audio intelligence LLM features out of the box"},{"t":"External LLM pipelines required","q":"requiring developers to orchestrate external LLM pipelines for deep analysis, summarization, or semantic extraction"}],"dropped":[{"t":"Nova models for conversational AI","q":"highly efficient Nova-2/Nova-3 models tailored for conversational AI agents"},{"t":"Customer data sharing required","q":"Standard pricing models require customer data sharing for model improvement"},{"t":"Privacy opt-out price premium","q":"forcing strict privacy-sensitive projects to pay a significant price premium to opt out"}]}],"api":"https://modelsagree.com/api/v1/best/best-speech-to-text-api.json"},{"slug":"best-speech-to-text-api-for-call-centers","title":"Best speech-to-text API for call centers","rank":1,"of":8,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Best overall balance of noisy 8 kHz telephony accuracy, low-latency streaming, diarization, keyterm prompting, redaction, scalability, and transparent pricing; narrowly leads AssemblyAI for live call-center workloads","reasons":[{"model":"ChatGPT","reason":"Best overall balance of noisy 8 kHz telephony accuracy, low-latency streaming, diarization, keyterm prompting, redaction, scalability, and transparent pricing; narrowly leads AssemblyAI for live call-center workloads"},{"model":"Claude","reason":"Purpose-built for exactly this workload — trained heavily on 8kHz telephony audio, sub-300ms streaming latency, multichannel and diarization out of the box, keyterm prompting for agent scripts/product names, and the best price-per-minute among top-tier providers at call-center volumes (~$0.0043-0.0077/min with volume discounts); also offers self-hosted/VPC deployment for compliance-bound contact centers. Ranked #1 on the assumption the typical practitioner needs real-time streaming plus batch at high volume, where its latency-accuracy-cost combination is unmatched"},{"model":"Gemini","reason":"Industry-leading speed and latency (sub-300ms streaming) optimized specifically for noisy, 8kHz telephony audio, under the assumption that low-latency real-time voice agents are the priority. Its conversational Flux models integrate native end-of-turn detection (EOT), simplifying the architecture of real-time voice agents by bypassing external Voice Activity Detection pipelines."},{"model":"Grok","reason":"Leading real-world accuracy"}],"fixes":[{"model":"ChatGPT","fix":"Advanced customization and private deployment generally require enterprise arrangements"},{"model":"Claude","fix":"Its built-in conversation intelligence (sentiment, topics, summarization) is thinner than AssemblyAI's — teams wanting rich analytics beyond the transcript must bolt on their own LLM layer"},{"model":"Gemini","fix":"A highly modular pricing model where essential contact center features like diarization, PII redaction, and formatting are billed as separate add-ons, which can quickly inflate production costs."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-speech-to-text-api-for-call-centers.json"},{"slug":"best-transcription-apis-for-real-time-voice-applications","title":"Best transcription APIs for real-time voice applications","rank":1,"of":8,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Best overall for English-first voice agents: strong Nova-3-class recognition plus model-integrated turn detection, interruption/resumption events, streaming partials, and roughly 260ms median end-of-turn detection; its eager endpoint events can start LLM work before a turn is fully committed.","reasons":[{"model":"ChatGPT","reason":"Best overall for English-first voice agents: strong Nova-3-class recognition plus model-integrated turn detection, interruption/resumption events, streaming partials, and roughly 260ms median end-of-turn detection; its eager endpoint events can start LLM work before a turn is fully committed."},{"model":"Claude","reason":"Purpose-built for streaming from day one — Nova-3 delivers ~300ms end-to-end latency over a real WebSocket API with word-level timestamps, endpointing, and interim results that actually behave predictably; pricing (~$0.0077/min streaming) undercuts most rivals at scale, and self-hosted/VPC deployment exists for regulated shops. For the typical voice-agent builder it has the best latency-accuracy-cost balance, which earns #1."},{"model":"Gemini","reason":"Industry-leading low latency, particularly with the Flux model featuring model-native turn detection under 260ms, and exceptionally high accuracy for conversational voice agents."},{"model":"Grok","reason":"Lowest real-world streaming latency (sub-300ms, p95 often under 300ms with integrated EOT detection saving 200-600ms on agent loops), strong accuracy on production voice agent audio (median ~6.8% WER), excellent for turn-taking in conversational apps, cost-effective at scale (~$0.005-0.008/min streaming), multilingual support with code-switching; assumes typical practitioner building interactive voice agents or live systems where speed and reliability trump niche features."}],"fixes":[{"model":"ChatGPT","fix":"Flux’s language coverage is narrower than the strongest multilingual APIs, and extracting maximum responsiveness requires tuning turn thresholds and handling speculative responses."},{"model":"Claude","fix":"Accuracy on heavily accented or far-field noisy audio still trails the best batch models, and its language coverage is thinner than Speechmatics or Google — English-first workloads shine, long-tail languages don't."},{"model":"Grok","fix":"Not the absolute top in entity preservation or broadest multilingual depth for highly code-switched/non-English heavy use cases (better for English-primary or specific languages)."}],"updated":"2026-07-18","api":"https://modelsagree.com/api/v1/best/best-transcription-apis-for-real-time-voice-applications.json"},{"slug":"best-speech-to-text-api-for-medical-transcription","title":"Best speech-to-text API for medical transcription","rank":2,"of":11,"score":13,"appearances":4,"modelRanks":{"ChatGPT":3,"Claude":2,"Gemini":4,"Grok":2},"reason":"Best developer-first option — a purpose-trained medical model with leading word-error and keyword recall on drug names and clinical terminology, real-time streaming latency suitable for live dictation, HIPAA-eligible with BAA, self-hosted deployment option, and dramatically lower per-minute cost than incumbents; near-tie with #1 for anyone building their own product rather than buying a clinician seat","reasons":[{"model":"Claude","reason":"Best developer-first option — a purpose-trained medical model with leading word-error and keyword recall on drug names and clinical terminology, real-time streaming latency suitable for live dictation, HIPAA-eligible with BAA, self-hosted deployment option, and dramatically lower per-minute cost than incumbents; near-tie with #1 for anyone building their own product rather than buying a clinician seat"},{"model":"Grok","reason":"Excellent low-latency streaming suitable for real-time clinical use, fine-tuned for medical vocabulary/pharma terms with strong real-world WER in noisy environments, BAA available, scalable pricing; high value for high-volume custom integrations serving practitioners."},{"model":"ChatGPT","reason":"Near-tied with AWS for API-first products thanks to fast streaming and batch transcription, strong medical-term recognition, diarization, practical developer tooling, and attractive latency-to-cost value"},{"model":"Gemini","reason":"Offers unmatched speed (under 300ms latency) and cost efficiency for raw medical transcription. It is highly optimized for complex medical terms, drugs, and accents, making it the premier choice for developers building highly responsive real-time voice agents or custom workflows."}],"fixes":[{"model":"ChatGPT","fix":"It produces transcription rather than a complete, clinically structured note workflow, so documentation generation and validation remain your responsibility"},{"model":"Claude","fix":"It's raw transcription — no built-in note structuring, EMR integrations, or clinician-facing workflow, so you build the documentation layer yourself"},{"model":"Gemini","fix":"Only provides raw text output, meaning developers must design, build, and maintain their own clinical summarization (SOAP note) and compliance pipelines."},{"model":"Grok","fix":"Slightly behind AssemblyAI on complex medical entity accuracy in head-to-head benchmarks; less bundled audio intelligence."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-speech-to-text-api-for-medical-transcription.json"},{"slug":"best-transcription-apis-for-speaker-diarization-in-meetings","title":"Best transcription APIs for speaker diarization in meetings","rank":2,"of":7,"score":7,"appearances":2,"modelRanks":{"Claude":3,"Gemini":2},"reason":"Industry-leading processing speed and ultra-low latency for both live-streaming and batch meeting transcription, paired with high diarization throughput at low per-hour costs. Near-tie with AssemblyAI for the top spot, assuming processing speed and cost efficiency at scale are the primary requirements.","reasons":[{"model":"Gemini","reason":"Industry-leading processing speed and ultra-low latency for both live-streaming and batch meeting transcription, paired with high diarization throughput at low per-hour costs. Near-tie with AssemblyAI for the top spot, assuming processing speed and cost efficiency at scale are the primary requirements."},{"model":"Claude","reason":"Fastest and among the cheapest at scale with genuinely usable diarization and true low-latency streaming, so it wins for live meeting captioning, high-volume pipelines, and cost-sensitive deployments; strong self-host/on-prem options for privacy-constrained buyers."}],"fixes":[{"model":"Claude","fix":"Diarization accuracy trails Speechmatics on hard multi-speaker audio, so it's not the pick when clean per-speaker attribution is the top priority over speed and price."},{"model":"Gemini","fix":"Diarization accuracy on noisy, heavily overlapping meeting audio slightly trails specialized diarization models without custom audio pre-processing."}],"updated":"2026-08-09","api":"https://modelsagree.com/api/v1/best/best-transcription-apis-for-speaker-diarization-in-meetings.json"},{"slug":"best-cheap-speech-to-text-api","title":"Best cheap speech-to-text API","rank":4,"of":10,"score":5,"appearances":2,"modelRanks":{"Claude":2,"Gemini":5},"reason":"Best accuracy-per-dollar among purpose-built STT vendors at ~$0.0043/min pay-as-you-go; Nova-3 beats Whisper-class models on noisy, multi-speaker, real-world audio, and you get streaming, diarization, keyterm prompting, and smart formatting in one API — the pick when the audio is hard, not just the price low","reasons":[{"model":"Claude","reason":"Best accuracy-per-dollar among purpose-built STT vendors at ~$0.0043/min pay-as-you-go; Nova-3 beats Whisper-class models on noisy, multi-speaker, real-world audio, and you get streaming, diarization, keyterm prompting, and smart formatting in one API — the pick when the audio is hard, not just the price low"},{"model":"Gemini","reason":"Optimized for real-time streaming conversational applications with pricing starting around $0.0043/min to $0.0052/min for batch, utilizing highly efficient Nova-2 and Nova-3 models."}],"fixes":[{"model":"Claude","fix":"5-6x Groq's price, so pure-batch users with clean audio overpay; volume discounts require committed-spend contracts"},{"model":"Gemini","fix":"Streaming endpoints are more expensive ($0.0077/min) and all intelligence features must be paid for as separate add-ons."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-cheap-speech-to-text-api.json"}],"page":"https://modelsagree.com/product/deepgram","check":"https://modelsagree.com/check?q=Deepgram","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}