Valkyrie by AzumoSchedule a call
Speech and audio trends · updated Oct 2, 2026

Which open-weight speech and audio models get downloaded most

A daily read of Hugging Face's top 5,000 models, focused on speech and audio models such as speech recognition and text to speech: the most downloaded, new uploads in the top 5,000, what they are used for, and who makes them.

Most downloaded

The most-downloaded open-weight speech and audio models

Ranked by downloads over the last 30 days. Fine-tuned versions count; most format conversions and re-uploads are left out. Counts include automated fetches, which lift models that tools and test suites pull by default. How models are placed is explained under About this data.

#ModelMakerTaskLicenceLast 30 daysDownloadsValkyrie
1wav2vec2-large-xlsr-53-japanesejonatasgrosmanSpeech recognitionApache-2.016.3MFine-tune one like it →Fine-tune one →
2Kokoro-82MhexgradText to speechApache-2.011.5MFine-tune one like it →Fine-tune one →
3clap-htsat-fusedlaionAudio classificationApache-2.07.3MFine-tune one like it →Fine-tune one →
4XTTS-v2coquiText to speechCustom licence6.7M
5whisper-large-v3-turboopenaiSpeech recognitionMIT6.4MFine-tune one like it →Fine-tune one →
6wav2vec2-large-xlsr-53-portuguesejonatasgrosmanSpeech recognitionApache-2.05.7MFine-tune one like it →Fine-tune one →
7speaker-diarization-community-1pyannoteSpeaker diarizationCC BY 4.05.5MFine-tune one like it →Fine-tune one →
8segmentation-3.0pyannoteVoice activity detectionMIT5.4MFine-tune one like it →Fine-tune one →
9wav2vec2-large-xlsr-53-russianjonatasgrosmanSpeech recognitionApache-2.04.3MFine-tune one like it →Fine-tune one →
10whisper-large-v3openaiSpeech recognitionApache-2.04.2MFine-tune one like it →Fine-tune one →
11whisper-smallopenaiSpeech recognitionApache-2.03.2MFine-tune one like it →Fine-tune one →
12mms-300m-1130-forced-alignerMahmoudAshrafSpeech recognitionCC BY-NC 4.02.8M
13Qwen3-TTS-12Hz-1.7B-CustomVoiceQwenText to speechApache-2.02.4MFine-tune one like it →Fine-tune one →
14wav2vec2-large-robust-24-ft-age-genderaudeeringAudio classificationCC BY-NC-SA 4.02.4M
15wav2vec2-large-xlsr-53-polishjonatasgrosmanSpeech recognitionApache-2.02.0MFine-tune one like it →Fine-tune one →
16musicgen-mediumfacebookText to audioCC BY-NC 4.01.9M
17Voxtral-Mini-4B-Realtime-2602mistralaiSpeech recognitionApache-2.01.9MFine-tune one like it →Fine-tune one →
18whisper-baseopenaiSpeech recognitionApache-2.01.8MFine-tune one like it →Fine-tune one →
19chatterboxResembleAIText to speechMIT1.7MFine-tune one like it →Fine-tune one →
20wav2vec2-base-960hfacebookSpeech recognitionApache-2.01.6MFine-tune one like it →Fine-tune one →
21Qwen3-ASR-1.7BQwenSpeech recognitionApache-2.01.6MFine-tune one like it →Fine-tune one →
22OmniVoicek2-fsaText to speech-1.4M
23whisper-tinyopenaiSpeech recognitionApache-2.01.4MFine-tune one like it →Fine-tune one →
24wav2vec2-large-xlsr-53-arabicjonatasgrosmanSpeech recognitionApache-2.01.4MFine-tune one like it →Fine-tune one →
25wav2vec2-large-xlsr-53-chinese-zh-cnjonatasgrosmanSpeech recognitionApache-2.01.4MFine-tune one like it →Fine-tune one →
New releases

New speech and audio models in the top 5,000

Uploaded to Hugging Face in the last six months with at least 25 likes or a trending score of 10, ranked by downloads over the last 30 days. Most format conversions and re-uploads are left out.

nemotron-3.5-asr-streaming-0.6bnvidiaSpeech recognition · uploaded May 15, 20261.2M downloads in 30 days
VieNeu-TTS-v3-Turbopnnbao-umpText to speech · uploaded Jun 5, 2026761K downloads in 30 daysFine-tune one like it →Fine-tune one →
sanoTTSampixaText to speech · uploaded Jul 13, 2026411K downloads in 30 days
MOSS-TTS-v1.5OpenMOSS-TeamText to speech · uploaded May 25, 2026130K downloads in 30 daysFine-tune one like it →Fine-tune one →
granite-speech-4.1-2bibm-graniteSpeech recognition · uploaded Apr 16, 2026122K downloads in 30 daysFine-tune one like it →Fine-tune one →
granite-speech-4.1-2b-plusibm-graniteSpeech recognition · uploaded Apr 16, 2026103K downloads in 30 daysFine-tune one like it →Fine-tune one →
What they're used for

Downloads by task

Share of downloads of speech and audio models in Hugging Face's top 5,000 on Oct 2, 2026, leaving out most format conversions and re-uploads.

Speech recognition58.5%
Text to speech20.8%
Audio classification9.3%
Voice activity detection4.1%
Speaker diarization3.7%
Audio to audio2.1%
Text to audio1.5%
Who's winning

Share of downloads by maker

Each repo counts under the account that uploaded it, leaving out most format conversions and re-uploads. Change is in percentage points since Jul 4, 2026.

jonatasgrosman22.6% (+5.7 pts)
openai11.1% (0.0 pts)
pyannote7.3% (+1.4 pts)
hexgrad7.1% (+0.8 pts)
Qwen4.8% (+1.5 pts)
laion4.6% (−1.4 pts)
facebook4.5% (+0.9 pts)
coqui4.2% (0.0 pts)
Other33.9% (−9.0 pts)

Licences of the most-downloaded speech and audio models

Apache-2.0 61.8%MIT 14.1%Llama licence 0.1%Other or none 24.1%

As declared on each repo, leaving out most format conversions and re-uploads.