Google DeepMind's Gemini 2.0 Launches Real-Time Multimodal Agent API with Live Audio and Visual Streams
Google has opened full access to the Gemini 2.0 Flash Multimodal Live API, enabling sub-200ms bi-directional voice and vision agent interactions for autonomous industrial applications.
# Gemini 2.0 Flash: Real-Time Multimodal Intelligence
Google DeepMind has commercially rolled out the Gemini 2.0 Flash Multimodal Live API. Supporting bi-directional WebRTC audio and video streaming, the model executes native reasoning over continuous camera feeds and audio inputs without intermediary transcription latencies.
The release also features the Google Search grounding tool and automated code execution sandbox, establishing Gemini 2.0 as a premier foundation for embodied agentic systems.
Primary Source Verified
Every claim in this report has been cross-referenced with official filings and documentation.