Google Releases Gemini 3.1 Flash Live: A Real-Time Multimodal Voice Model for Low-Latency Audio, Video, and Tool Use for AI Agents
By Asif Razzaq
As covered in News yesterday, Google released Gemini 3.1 Flash Live in preview, a real-time multimodal voice model designed for low-latency audio, video, and tool use. It natively processes multimodal streams, eliminating the traditional 'wait-time stack' of turn-based LLM voice architectures.