• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Bengali word error rate is claimed to have fallen from 85% to 7.65% after one fine-tune

    The poster says 80 labs asked to license Monsoon in a week and praises VoiceArena’s approach to building datasets.

    elvisEL
    Rohan PaulRP
    2 Sources, ,

    TLDR

    In an October 8 post, a user claimed one fine-tune brought Bengali word error rate (WER) down from 85% to 7.65%. They also said 80 labs asked to license Monsoon in a week and praised what they described as VoiceArena’s rule to build datasets only if they “move the needle.”

    Combined views

    5.1K

    2 Sources, first seen 2h ago

    Combined views

    5.1K

    2 Sources, first seen 2h ago

    24 likes
    2h ago
    first seen 2h ago
    24 likes
    10 comments
    8 saves
    2 reposts
    10 comments
    8 saves
    2 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    elvis@omarsar0WER dropped from 85% to 7.65% on Bengali with one fine-tune. Numbers like that are why 80 labs asked to license Monsoon in a week. What I respect most is the rule behind it: @voicearena_ai only builds a dataset if it moves the needle. Most data vendors can't say that.2h
    Rohan Paul@rohanpaul_aiFor ASR (automatic speech recognition) on weak languages, more training data tends to beat a bigger model. @voicearena_ai’s Bengali result for Monsoon, its ASR corpus, is shows it. Whisper Medium fine-tuned on Monsoon went from 85.27% to 7.65% LLM word error rate on Bengali FLEURS. Medium is the 769M-parameter Whisper. Small enough to serve cheaply, and in many settings small enough to run close to the user.2h

    2 Sources

    elvis@omarsar0WER dropped from 85% to 7.65% on Bengali with one fine-tune. Numbers like that are why 80 labs asked to license Monsoon in a week. What I respect most is the rule behind it: @voicearena_ai only builds a dataset if it moves the needle. Most data vendors can't say that.2h
    Rohan Paul@rohanpaul_aiFor ASR (automatic speech recognition) on weak languages, more training data tends to beat a bigger model. @voicearena_ai’s Bengali result for Monsoon, its ASR corpus, is shows it. Whisper Medium fine-tuned on Monsoon went from 85.27% to 7.65% LLM word error rate on Bengali FLEURS. Medium is the 769M-parameter Whisper. Small enough to serve cheaply, and in many settings small enough to run close to the user.2h