NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition
A user says the model never rewrites committed text and can use names, product terms and jargon supplied at runtime to steer recognition.
TLDR
A user says NetEase Youdao has open-sourced Confucius4-R2T2, a streaming speech-recognition model built on Qwen3-ASR. The post describes its Longest Stable Prefix learning as deciding when text is safe to emit and when more audio is needed, while leaving committed text unchanged. It also says the LLM-based decoder can use context supplied at runtime without changing model weights. The user argues these features matter for voice agents that might act on transcripts that later change, and for meeting or coding assistants that need to recognize specialized vocabulary.
Combined views
193
1 Source, first seen 1d ago
NetEase Youdao open-sources Confucius4-R2T2 for streaming speech recognition
A user says the model never rewrites committed text and can use names, product terms and jargon supplied at runtime to steer recognition.
TLDR
A user says NetEase Youdao has open-sourced Confucius4-R2T2, a streaming speech-recognition model built on Qwen3-ASR. The post describes its Longest Stable Prefix learning as deciding when text is safe to emit and when more audio is needed, while leaving committed text unchanged. It also says the LLM-based decoder can use context supplied at runtime without changing model weights. The user argues these features matter for voice agents that might act on transcripts that later change, and for meeting or coding assistants that need to recognize specialized vocabulary.