September 16, 2026

Ⓜ️ Google’s new voice models just topped the speech-to-speech leaderboard

Article featured image

Ⓜ️ Google’s new voice models just topped the speech-to-speech leaderboard Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, targeting real-time voice agents rather than traditional text-based interactions. Gemini 3.8 Live reportedly reaches frontier-level performance at around $0.84 per hour, while delivering a higher composite score than GPT-Realtime-2 High at roughly 80% lower measured cost. Both models are multimodal: voice is the main interface, but they can also process visual information during a conversation and automatically switch between 97 supported languages. The biggest difference comes with Extended Thinking. It scores 68.6% on τ-Voice, narrowly ahead of GPT-Live-1 Astra Medium at 67.9%. τ-Voice isn’t simply testing whether an AI sounds natural. It measures whether a voice agent can actually complete complex, multi-step customer-service tasks while following policies, using tools correctly and reaching the right outcome across airline, retail and telecom scenarios. @aipost 🏴

Article image 2