• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition

    A user says the model never rewrites committed text and can use names, product terms and jargon supplied at runtime to steer recognition.

    Rohan PaulRP
    1 Source, 21d ago, first seen 21d ago

    TLDR

    A user says NetEase Youdao has open-sourced Confucius4-R2T2, a streaming speech-recognition model built on Qwen3-ASR. The post describes its Longest Stable Prefix learning as deciding when text is safe to emit and when more audio is needed, while leaving committed text unchanged. It also says the LLM-based decoder can use context supplied at runtime without changing model weights. The user argues these features matter for voice agents that might act on transcripts that later change, and for meeting or coding assistants that need to recognize specialized vocabulary.

    Combined views

    193

    1 Source, first seen 21d ago

    Combined views

    193

    1 Source, first seen 21d ago

    1 likes
    1 likes
    2 comments
    2 comments

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Rohan Paul@rohanpaul_aiNow, why thats so important is that generic ASR (Automatic Speech Recognition) collapses exactly where real work happens. A meeting assistant that can't spell your company's internal tool names is not a meeting assistant. A coding voice agent that mangles technical vocabulary is just not usable. The failure mode isn't a wrong word here and there. It's that the system quietly becomes unusable in the one setting you built it for.21d

    1 Source

    Rohan Paul@rohanpaul_aiNow, why thats so important is that generic ASR (Automatic Speech Recognition) collapses exactly where real work happens. A meeting assistant that can't spell your company's internal tool names is not a meeting assistant. A coding voice agent that mangles technical vocabulary is just not usable. The failure mode isn't a wrong word here and there. It's that the system quietly becomes unusable in the one setting you built it for.21d