Parakeet Redux launches as a 178MB speech-to-text model
Its developer says it compresses Nvidia’s Parakeet from 1.2GB to 178MB, runs at 113 times real time on CPU and beats the base model on the 25-language FLEURS benchmark.
TLDR
Parakeet Redux’s developer announced the release of a ternary speech-to-text model compressed from Nvidia’s Parakeet. They claim the 178MB model runs at 113 times real time on CPU and beats the base model on the 25-language FLEURS benchmark while staying within 0.3 WER (word error rate) on English. The developer says their quantization recipe minimizes forgetting when moving from 16-bit to ternary, and that post-training recovers most of the lost accuracy. They also shared a Hugging Face model card described as containing full benchmark tables and usage examples.
Parakeet Redux launches as a 178MB speech-to-text model
Its developer says it compresses Nvidia’s Parakeet from 1.2GB to 178MB, runs at 113 times real time on CPU and beats the base model on the 25-language FLEURS benchmark.
