• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition

    A post describing the release says the model never rewrites committed text, targeting a risk for voice agents: acting on partial transcripts that later change.

    Rohan PaulRP
    2 Sources, 21d ago, first seen 21d ago

    TLDR

    A post describes NetEase Youdao’s Confucius4-R2T2 as an open-source streaming speech-recognition model built on Qwen3-ASR. It says the model uses Longest Stable Prefix learning to decide when text is safe to emit and when more audio is needed, without rewriting committed text. According to the post, its language-model-based decoder also accepts context at runtime—such as names, product terms, industry jargon and meeting topics—to steer recognition without changing the model’s weights.

    Combined views

    6.1K

    2 Sources, first seen 21d ago

    Combined views

    6.1K

    2 Sources, first seen 21d ago

    19 likes
    19 likes
    9 comments
    7 saves
    12 reposts
    9 comments
    7 saves
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Rohan Paul@rohanpaul_aiVoice agents should consume speech incrementally but only act on committed text, because a fast transcript that mutates text can corrupt downstream agent state. NetEase Youdao just open-sourced Confucius4-R2T2, a streaming ASR (Automatic Speech Recognition) model built exactly around that constraint. It never rewrites committed text, i.e. my text is never gets rewritten underneath me. That append-only behavior targets a very serioius production failure in voice agents, where software may act on partial speech before the speaker finishes. Built on Qwen3-ASR, R2T2 uses Longest Stable Prefix learning to decide when text is safe to emit and when it needs more audio context. The underrated detail in Confucius R2T2 is that because the decoder is LLM-based, context can be injected at runtime. Names. Product terms. Industry jargon. Meeting topics. You steer recognition without touching the weights, which is a very different design choice from treating the acoustic model as a fixed black box.21d

    2 Sources

    Rohan Paul@rohanpaul_aiVoice agents should consume speech incrementally but only act on committed text, because a fast transcript that mutates text can corrupt downstream agent state. NetEase Youdao just open-sourced Confucius4-R2T2, a streaming ASR (Automatic Speech Recognition) model built exactly around that constraint. It never rewrites committed text, i.e. my text is never gets rewritten underneath me. That append-only behavior targets a very serioius production failure in voice agents, where software may act on partial speech before the speaker finishes. Built on Qwen3-ASR, R2T2 uses Longest Stable Prefix learning to decide when text is safe to emit and when it needs more audio context. The underrated detail in Confucius R2T2 is that because the decoder is LLM-based, context can be injected at runtime. Names. Product terms. Industry jargon. Meeting topics. You steer recognition without touching the weights, which is a very different design choice from treating the acoustic model as a fixed black box.21d