Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
Google released native speech-to-speech models via Gemini API emphasizing cost-efficiency, fluid dialogue, visual grounding, and background tool calls. Extended Thinking variant adds multi-step reasoning and tops benchmarks like Artificial Analysis' Speech-to-Speech Quality Index.
TLDR
Advances practical voice agents that feel natural and handle complexity without breaking flow—key for consumer and enterprise use. Strong benchmarks and lower costs fuel adoption comparisons; ties into narrative of AI doing the job, not just answering. Voice emerging as interface for agents.
Combined views
—
2 Sources, first seen 1d ago
Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
Google released native speech-to-speech models via Gemini API emphasizing cost-efficiency, fluid dialogue, visual grounding, and background tool calls. Extended Thinking variant adds multi-step reasoning and tops benchmarks like Artificial Analysis' Speech-to-Speech Quality Index.
TLDR
Advances practical voice agents that feel natural and handle complexity without breaking flow—key for consumer and enterprise use. Strong benchmarks and lower costs fuel adoption comparisons; ties into narrative of AI doing the job, not just answering. Voice emerging as interface for agents.