Report
PocketTTS training blog details a ‘drifting’ approach to on-device speech synthesis
Kyutai says its 100-million-parameter model achieves less than 1% word error rate and supports voice cloning.
TLDR
Kyutai released a technical blog on training PocketTTS, its on-device text-to-speech model, using drifting, a one-step generative objective from Deng et al. The lab reports less than 1% word error rate and voice cloning. To its knowledge, PocketTTS is the first speech model and the first autoregressive model trained this way.
Combined views
4.1K
2 Sources, first seen ago
