ModelsAgree
← All leaderboards

NVIDIA Canary-Qwen 2.5B

What ChatGPT, Claude, Gemini & Grok actually say · August 2026

Visit nvidia.com

The verdict

NVIDIA Canary-Qwen 2.5B appears in 1 AI-ranked category — best position #3 for open-source speech-to-text model.

Positioning brief — for the NVIDIA Canary-Qwen 2.5B team

Why the models put NVIDIA Canary-Qwen 2.5B at #3 for open-source speech-to-text model

  • Exceptional English accuracy Grok · Claude · GPTExceptional English accuracy, punctuation, capitalization, hallucination robustness
  • hybrid ASR+LLM design Grok · Claudea hybrid ASR+LLM design that also enables post-transcription tasks
  • summarization or question answering Claude · GPToptional transcript summarization or question answering

What the models credit OpenAI Whisper (#1) with — and don’t credit NVIDIA Canary-Qwen 2.5B

  • multilingual coverage across 99+ languages Claude · Gemini · GPT · GrokUnmatched multilingual coverage across 99+ languages
  • robustness to accents and noisy audio Claude · GPT · Grokstrong robustness to accents and noisy real-world audio
  • runs everywhere from a phone to server Claude · GPTit runs everywhere from a phone to a server

What would move the rank — the models’ fix lines, unified

  • English-only GPT · ClaudeEnglish-only
  • Higher latency and compute needs GrokHigher latency/compute needs than lighter CTC/TDT models
  • not ideal for real-time streaming Groknot ideal for real-time streaming or very low-resource/edge devices without optimization

Restructured from verbatim model output · nothing invented · every quote machine-verified

#3🗣 Best open-source speech-to-text model3/4 models · updated 2026-07-15
GPT #5Claude #4Gemini Grok #1

Tops or near-tops Hugging Face Open ASR Leaderboard with ~5.63% avg WER on English benchmarks via strong Conformer + Qwen LLM decoder hybrid (SALM architecture); excellent accuracy for typical practitioner use in clean-to-moderate noisy English audio; strong community/HF support and NVIDIA tooling for deployment.

Claude Took the top WER spot on the Open ASR Leaderboard at release with a hybrid ASR+LLM design that also enables post-transcription tasks (summarization, Q&A), CC-BY licensed and commercially usable — near-tie with Parakeet, which wins above it only on throughput per dollar for bulk transcription.

GPT Exceptional English accuracy, punctuation, capitalization, hallucination robustness, and optional transcript summarization or question answering

Where NVIDIA Canary-Qwen 2.5B falls short, per the models

  • GPT It is English-only, relatively large, NeMo-dependent, and optimized around audio segments no longer than roughly 40 seconds
  • Claude English-only, which disqualifies it for the many practitioners who need multilingual support (its multilingual Canary 1B sibling covers fewer languages at lower accuracy).
  • Grok Higher latency/compute needs than lighter CTC/TDT models; not ideal for real-time streaming or very low-resource/edge devices without optimization.

Top alternatives per the models: OpenAI Whisper · NVIDIA Parakeet TDT · IBM Granite Speech · Qwen3-ASR

Head-to-head — how the models call it

Watch NVIDIA Canary-Qwen 2.5B

Boards re-poll weekly and the models change their minds. One short email only when NVIDIA Canary-Qwen 2.5B's standing moves — a rank change, a rival overtaking, or new reasoning from the models. Nothing otherwise.

Embed your ranking badge

NVIDIA Canary-Qwen 2.5B ranks #3 for best open-source speech-to-text model by AI-model consensus. Put the badge in your README, docs or site — it updates automatically as the models re-rank.

NVIDIA Canary-Qwen 2.5B — ranked #3 for Best open-source speech-to-text model by AI models on ModelsAgree
Markdown (README)
[![NVIDIA Canary-Qwen 2.5B — ranked #3 for Best open-source speech-to-text model by AI models on ModelsAgree](https://modelsagree.com/badge/nvidia-canary-qwen-2-5b.svg)](https://modelsagree.com/best/best-open-source-speech-to-text-model?utm_source=badge&utm_medium=embed&utm_campaign=badge-nvidia-canary-qwen-2-5b)
HTML
<a href="https://modelsagree.com/best/best-open-source-speech-to-text-model?utm_source=badge&utm_medium=embed&utm_campaign=badge-nvidia-canary-qwen-2-5b"><img src="https://modelsagree.com/badge/nvidia-canary-qwen-2-5b.svg" alt="NVIDIA Canary-Qwen 2.5B — ranked #3 for Best open-source speech-to-text model by AI models on ModelsAgree" height="28"></a>

Rankings are computed from what the models answer, re-polled on demand · raw reasoning shown verbatim · methodology