R2T2 is said to transcribe speech as you speak
A user describes R2T2 as a speech-to-text model that processes audio in small chunks, publishing words it considers stable while holding back uncertain words for more context.
TLDR
A user shares a Hugging Face link for R2T2, describing a model that turns speech into text as it comes in. In a follow-up reply, they explain that it publishes words it considers stable, holds back uncertain words and releases more text as it gathers context.
Combined views
8.6K
2 Sources, first seen 6h ago
R2T2 is said to transcribe speech as you speak
A user describes R2T2 as a speech-to-text model that processes audio in small chunks, publishing words it considers stable while holding back uncertain words for more context.
TLDR
A user shares a Hugging Face link for R2T2, describing a model that turns speech into text as it comes in. In a follow-up reply, they explain that it publishes words it considers stable, holds back uncertain words and releases more text as it gathers context.