{"slug":"deepinfra","name":"DeepInfra","domain":"deepinfra.com","verdict":"As of 2026-07-15, ChatGPT, Claude, Gemini, Grok collectively rank DeepInfra #4 of 8 for serverless llm inference api (one of 2 leaderboards it appears on). Source: https://modelsagree.com/product/deepinfra (modelsagree.com, CC BY 4.0).","best_rank":4,"categories":2,"entries":[{"slug":"best-serverless-llm-inference-api","title":"Best serverless LLM inference API","rank":4,"of":8,"score":9,"appearances":4,"modelRanks":{"ChatGPT":4,"Claude":4,"Gemini":3,"Grok":4},"reason":"The value champion offering the lowest pricing for standard open-weights models (like Llama and Qwen) while maintaining excellent throughput. Near-tied with Fireworks AI on speed metrics, but ranked lower due to fewer developer features.","reasons":[{"model":"Gemini","reason":"The value champion offering the lowest pricing for standard open-weights models (like Llama and Qwen) while maintaining excellent throughput. Near-tied with Fireworks AI on speed metrics, but ranked lower due to fewer developer features."},{"model":"ChatGPT","reason":"Excellent price-to-performance, broad open-model selection, simple pay-as-you-go OpenAI-compatible APIs, embeddings and rerankers, and low operational friction make it especially compelling for cost-conscious applications."},{"model":"Claude","reason":"Consistently the lowest per-token prices on a wide open-model catalog with a no-friction OpenAI-compatible API — the best value pick for cost-sensitive, high-volume workloads like batch processing, embeddings, and classification."},{"model":"Grok","reason":"best cost-efficiency with low per-token rates, reliable serverless scaling, solid model selection for budget-conscious production workloads"}],"fixes":[{"model":"ChatGPT","fix":"It offers fewer production-control, capacity-guarantee, and enterprise deployment options than the top two."},{"model":"Claude","fix":"Latency and throughput are less consistent than Fireworks/Together and the enterprise feature set (compliance, SLAs, dedicated capacity options) is thinner — not for latency-sensitive production frontends."},{"model":"Gemini","fix":"Lacks advanced agentic tooling, structured output optimizations, or robust fine-tuning options, making it unsuitable for highly customized agent pipelines."},{"model":"Grok","fix":"improve raw speed and TTFT to match leaders for latency-sensitive use cases"}],"updated":"2026-07-15","rank_history":{"days":["2026-06-29","2026-06-30","2026-07-08","2026-07-09","2026-07-10","2026-07-12","2026-07-13","2026-07-14","2026-07-15"],"ranks":[4,null,7,4,null,4,4,4,4]},"reasoning_shift":[{"model":"Gemini","from":"2026-07-14","to":"2026-07-15","added":[{"t":"excellent throughput","q":"while maintaining excellent throughput"},{"t":"near-tied on speed metrics","q":"Near-tied with Fireworks AI on speed metrics"},{"t":"agentic tooling and structured outputs","q":"Lacks advanced agentic tooling, structured output optimizations"}],"dropped":[{"t":"OpenAI-compatible API","q":"a reliable OpenAI-compatible API"},{"t":"virtually no cold-start latency","q":"virtually no cold-start latency"},{"t":"serverless custom LoRA deployments","q":"serverless custom LoRA deployments"}]},{"model":"ChatGPT","from":"2026-07-14","to":"2026-07-15","added":[{"t":"low operational friction","q":"low operational friction"},{"t":"capacity-guarantee options","q":"capacity-guarantee"},{"t":"enterprise deployment options","q":"enterprise deployment options"}],"dropped":[{"t":"vision and audio coverage","q":"vision, and audio coverage"},{"t":"performance consistency","q":"performance consistency"},{"t":"documentation and support","q":"documentation, and support are less dependable than the top three"}]},{"model":"Claude","from":"2026-07-13","to":"2026-07-14","added":[{"t":"batch processing, embeddings, and classification","q":"batch processing, embeddings, and classification"},{"t":"latency and throughput less consistent","q":"Latency and throughput are less consistent than Fireworks/Together"},{"t":"not for latency-sensitive production frontends","q":"not for latency-sensitive production frontends"}],"dropped":[{"t":"thinner support and tooling polish","q":"Thinner support, SLAs, and tooling polish than Together/Fireworks"}]}],"api":"https://modelsagree.com/api/v1/best/best-serverless-llm-inference-api.json"},{"slug":"best-cheap-speech-to-text-api","title":"Best cheap speech-to-text API","rank":5,"of":10,"score":5,"appearances":1,"modelRanks":{"Gemini":1},"reason":"Offers the lowest raw pricing on the market at $0.00020/min for Whisper Large V3 Turbo and $0.00045/min for Whisper Large V3, running on high-concurrency serverless GPU infrastructure.","reasons":[{"model":"Gemini","reason":"Offers the lowest raw pricing on the market at $0.00020/min for Whisper Large V3 Turbo and $0.00045/min for Whisper Large V3, running on high-concurrency serverless GPU infrastructure."}],"fixes":[{"model":"Gemini","fix":"Provides only raw transcription outputs without advanced features like diarization, custom vocabulary, or audio intelligence."}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-cheap-speech-to-text-api.json"}],"page":"https://modelsagree.com/product/deepinfra","check":"https://modelsagree.com/check?q=DeepInfra","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}