Routing voice AI tool calls through a text backend
A post discussing NVIDIA research reports a gap on grounded customer-service tasks under clean conditions: commercial duplex voice models complete 31–51%, versus 85% for text agents such as GPT-5 on the same tasks in text mode.
TLDR
A post describes an NVIDIA research architecture that delegates tool decisions to a text model rather than the speech model. The full-duplex speech frontend streams transcripts to the text backend, which handles the tool call; the result is passed back for text-to-speech output. The post reports 92.0–97.2% tool-call recall and 81.2% accuracy at rejecting irrelevant calls. It says turn-taking rate, streaming speech-recognition word error rate and spoken-language intelligence remain at the no-tool-call baseline.
Combined views
7K
1 Source, first seen 11h ago
Routing voice AI tool calls through a text backend
A post discussing NVIDIA research reports a gap on grounded customer-service tasks under clean conditions: commercial duplex voice models complete 31–51%, versus 85% for text agents such as GPT-5 on the same tasks in text mode.