Ask a question below.
Published answers will appear here.
does your audio-input LLM secretly know how to generate speech? yes but it sounds like a scary demon! deepdream but for speech: starting with noise audio, I tried optimizing the probability of a desired transcription, while keeping the model frozen. volume up! 🎧
this is Gemma 4 e4b, with some auxiliary losses from whisper and wav2vec. Gemma on its own sort of works, but it tends to repeat words a bit more and is less clear.
Ask a question below.
Published answers will appear here.