WaveNet co-creator traces a decade of speech AI
The researcher describes a path from generating raw audio waveforms to multimodal systems such as Gemini, and anticipates more expressive speech, faster dialogue and integration with robots.
TLDR
In a September 8, 2026 retrospective, a researcher who helped create WaveNet marks 10 years since announcing the speech synthesis system. Their account traces developments from WaveNet’s direct generation of raw audio waveforms through systems that encode audio as tokens for language models, then to multimodal models such as Gemini that combine different types of information. Looking toward the next decade, they expect finer expression of context and emotion, very low-latency dialogue and integration with robots.
Combined views
11.1K
3 Sources, first seen 22d ago