Voice AI’s challenge of keeping up with a growing conversation
A post argues that getting the answer right isn’t enough for voice AI: models also have to keep up as conversations get longer. It suggests continuous state may be a better way to frame the problem than speech.
TLDR
A post questions whether the same model architecture should work equally well for thinking and interacting. It describes a video of Cartesia’s founder discussing the company’s roots in sequence modeling rather than speech research, and speculates that this background explains its focus on state space models. The author links to the full video on The Neon Show’s YouTube channel.