Announcement
Kyutai releases technical blog on training PocketTTS with drifting
Kyutai says its 100-million-parameter, on-device text-to-speech model has a word error rate below 1% and supports voice cloning.
TLDR
Kyutai released a technical blog on training PocketTTS with drifting, a recent one-step generative objective from Deng et al. The lab claims a word error rate below 1%, high-quality speech and voice cloning. To its knowledge, PocketTTS is the first speech model and the first autoregressive model trained this way.
Combined views
7.5K
2 Sources, first seen 2h ago
