Monsoon launches with unscripted speech data in 50 languages
A post describes VoiceArena’s dataset as conversations from 23 countries, with short, noisy clips kept in the training mix rather than filtered out.
TLDR
A post says VoiceArena fine-tuned Whisper Medium on Monsoon and measured a 16.1% semantic word error rate on IndicVoices Telugu, compared with 92.7% for the off-the-shelf model. The author attributes the 76.6-point drop to the training data, not a larger model or a change in architecture.
Combined views
5.4K
3 Sources, first seen 14h ago
