• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Routing voice AI tool calls through a text backend

A post discussing NVIDIA research reports a gap on grounded customer-service tasks under clean conditions: commercial duplex voice models complete 31–51%, versus 85% for text agents such as GPT-5 on the same tasks in text mode.

elvisEL
2 Sources, 20d ago, first seen 20d ago

TLDR

A post describes an NVIDIA research architecture that delegates tool decisions to a text model rather than the speech model. The full-duplex speech frontend streams transcripts to the text backend, which handles the tool call; the result is passed back for text-to-speech output. The post reports 92.0–97.2% tool-call recall and 81.2% accuracy at rejecting irrelevant calls. It says turn-taking rate, streaming speech-recognition word error rate and spoken-language intelligence remain at the no-tool-call baseline.

Combined views

10.9K

2 Sources, first seen 20d ago

96 likes22 comments101 saves27 reposts

Combined views

10.9K

2 Sources, first seen 20d ago

96 likes22 comments101 saves27 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

2 Sources

elvis@omarsar0NVIDIA research papers are on fire recently! Here is another interesting paper where they give full-duplex speech models tool calls. (bookmark it) Commercial duplex voice models complete 31 to 51 percent of grounded customer-service tasks under clean conditions. Text agents like GPT-5 reach 85 percent on the same tasks in text mode. Most of what a voice agent loses, it loses in the speech pipeline. The fix routes the decision out of the speech model. The duplex frontend learns to emit a delegation token, forwards streaming transcripts to a text backend LLM for the tool call, and receives the result through a lightweight prefill-and-repeat mechanism before streaming TTS speaks it. Tool-call recall runs 92.0 to 97.2 percent with 81.2 percent accuracy at rejecting irrelevant calls. Turn-taking rate, streaming ASR word error rate and spoken-language intelligence all stay at the no-tool-call baseline. Paper: https://arxiv.org/abs/2609.19334 Chat with Paper: https://academy.dair.ai/papers/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models-2609.1933420d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    2 Sources

    elvis@omarsar0NVIDIA research papers are on fire recently! Here is another interesting paper where they give full-duplex speech models tool calls. (bookmark it) Commercial duplex voice models complete 31 to 51 percent of grounded customer-service tasks under clean conditions. Text agents like GPT-5 reach 85 percent on the same tasks in text mode. Most of what a voice agent loses, it loses in the speech pipeline. The fix routes the decision out of the speech model. The duplex frontend learns to emit a delegation token, forwards streaming transcripts to a text backend LLM for the tool call, and receives the result through a lightweight prefill-and-repeat mechanism before streaming TTS speaks it. Tool-call recall runs 92.0 to 97.2 percent with 81.2 percent accuracy at rejecting irrelevant calls. Turn-taking rate, streaming ASR word error rate and spoken-language intelligence all stay at the no-tool-call baseline. Paper: https://arxiv.org/abs/2609.19334 Chat with Paper: https://academy.dair.ai/papers/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models-2609.1933420d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet