NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition
A post describing the release says the model never rewrites committed text, targeting a risk for voice agents: acting on partial transcripts that later change.
TLDR
A post describes NetEase Youdao’s Confucius4-R2T2 as an open-source streaming speech-recognition model built on Qwen3-ASR. It says the model uses Longest Stable Prefix learning to decide when text is safe to emit and when more audio is needed, without rewriting committed text. According to the post, its language-model-based decoder also accepts context at runtime—such as names, product terms, industry jargon and meeting topics—to steer recognition without changing the model’s weights.
Combined views
5.5K
2 Sources, first seen 1d ago
NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition
A post describing the release says the model never rewrites committed text, targeting a risk for voice agents: acting on partial transcripts that later change.
TLDR
A post describes NetEase Youdao’s Confucius4-R2T2 as an open-source streaming speech-recognition model built on Qwen3-ASR. It says the model uses Longest Stable Prefix learning to decide when text is safe to emit and when more audio is needed, without rewriting committed text. According to the post, its language-model-based decoder also accepts context at runtime—such as names, product terms, industry jargon and meeting topics—to steer recognition without changing the model’s weights.