{"slug":"groq","name":"Groq","domain":"groq.com","verdict":"As of 2026-07-15, ChatGPT, Claude, Gemini, Grok collectively rank Groq first for cheap speech-to-text api (one of 2 leaderboards it appears on). Source: https://modelsagree.com/product/groq (modelsagree.com, CC BY 4.0).","best_rank":1,"categories":2,"brief":{"category":"best-cheap-speech-to-text-api","title":"Best cheap speech-to-text API","rank":1,"of":10,"top":null,"day":"2026-07-16","why":[{"t":"accurate multilingual transcription","m":["Claude","ChatGPT"],"q":"accurate multilingual transcription at $0.04/hour"},{"t":"extremely fast inference","m":["Claude","ChatGPT","Gemini"],"q":"extremely fast inference"},{"t":"nothing hosted comes close on price","m":["Claude","ChatGPT","Gemini"],"q":"nothing hosted comes close on price-per-accurate-minute"},{"t":"easy OpenAI-compatible API","m":["ChatGPT"],"q":"an easy OpenAI-compatible API make it unusually practical"}],"gap":[],"fix":[{"t":"no streaming","m":["ChatGPT","Claude"],"q":"no streaming"},{"t":"bare-bones feature set","m":["ChatGPT","Claude"],"q":"bare-bones feature set"},{"t":"strict rate limits","m":["Claude","Gemini"],"q":"strict rate limits on request concurrency and volume"}]},"entries":[{"slug":"best-cheap-speech-to-text-api","title":"Best cheap speech-to-text API","rank":1,"of":10,"score":12,"appearances":3,"modelRanks":{"ChatGPT":2,"Claude":1,"Gemini":3},"reason":"Cheapest hosted transcription that's still genuinely accurate — roughly $0.04/hour (~$0.0007/min) for OpenAI's Whisper large-v3-turbo weights served on Groq's LPUs, with faster-than-real-time turnaround measured in seconds; for the typical developer batch-transcribing podcasts, calls, or user audio, nothing hosted comes close on price-per-accurate-minute","reasons":[{"model":"Claude","reason":"Cheapest hosted transcription that's still genuinely accurate — roughly $0.04/hour (~$0.0007/min) for OpenAI's Whisper large-v3-turbo weights served on Groq's LPUs, with faster-than-real-time turnaround measured in seconds; for the typical developer batch-transcribing podcasts, calls, or user audio, nothing hosted comes close on price-per-accurate-minute"},{"model":"ChatGPT","reason":"Near-tied for first: accurate multilingual transcription at $0.04/hour, extremely fast inference, and an easy OpenAI-compatible API make it unusually practical."},{"model":"Gemini","reason":"Delivers near-instantaneous transcription speed (up to 200x real-time) at $0.00067/min for Whisper Large V3 Turbo, using highly optimized LPU hardware."}],"fixes":[{"model":"ChatGPT","fix":"File transcription rather than true streaming, with minimum-duration billing and fewer speech-specific features than specialist APIs."},{"model":"Claude","fix":"Batch/pre-recorded only with a bare-bones feature set — no streaming, weak diarization, and rate limits on lower tiers make it wrong for real-time voice agents or enterprise SLAs"},{"model":"Gemini","fix":"Subject to strict rate limits on request concurrency and volume, making it difficult to scale without enterprise negotiation."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-cheap-speech-to-text-api.json"},{"slug":"best-serverless-llm-inference-api","title":"Best serverless LLM inference API","rank":3,"of":8,"score":13,"appearances":4,"modelRanks":{"ChatGPT":3,"Claude":3,"Gemini":4,"Grok":1},"reason":"unmatched ultra-low latency and highest tokens/sec (500-1000+ tps on LPUs) with no cold starts for real-time apps like agents/voice","reasons":[{"model":"Grok","reason":"unmatched ultra-low latency and highest tokens/sec (500-1000+ tps on LPUs) with no cold starts for real-time apps like agents/voice"},{"model":"ChatGPT","reason":"Exceptional token-generation speed and low latency on supported models, straightforward OpenAI-compatible integration, competitive token pricing, and discounted batch processing make it the strongest choice for latency-sensitive interactive applications."},{"model":"Claude","reason":"LPU hardware delivers hundreds of tokens/sec at low per-token prices, making it the clear winner for latency-critical UX (voice agents, real-time chat, agent loops) where time-to-final-token dominates the experience."},{"model":"Gemini","reason":"Delivers unmatched, class-leading inference speeds and sub-second Time to First Token (TTFT) using custom LPU hardware, making it essential for conversational voice applications."}],"fixes":[{"model":"ChatGPT","fix":"The curated model catalog and hardware-specific support are substantially narrower than Fireworks or Together, limiting model choice and customization."},{"model":"Claude","fix":"A narrow, slow-moving model menu with no custom-model or fine-tuned deployment support, plus capacity/rate-limit constraints at scale — unusable if your model isn't on their list."},{"model":"Gemini","fix":"High rate limits and restricted context windows make it unsuitable for high-volume document processing or long-context RAG applications."},{"model":"Grok","fix":"dramatically expand model catalog beyond narrow selection of optimized models"}],"updated":"2026-07-15","rank_history":{"days":["2026-06-29","2026-06-30","2026-07-08","2026-07-09","2026-07-10","2026-07-12","2026-07-13","2026-07-14","2026-07-15"],"ranks":[3,1,4,3,3,3,3,3,3]},"api":"https://modelsagree.com/api/v1/best/best-serverless-llm-inference-api.json"}],"page":"https://modelsagree.com/product/groq","check":"https://modelsagree.com/check?q=Groq","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}