Google unveils Gemini Omni, a multimodal AI model that generates video from text, images, and audio The multimodal model turns text, images, audio, and existing footage into realistic video clips, with implications that ripple well beyond Mountain View. Google DeepMind just dropped what might be the most capable video generation model yet. Gemini Omni, unveiled at Google I/O on May 19-20, 2026, accepts text, images, audio, and video as inputs and spits out short video clips, roughly 10 seconds long, complete with synchronized audio. The model’s first variant, Gemini Omni Flash, is the tip of the spear. It replaces Google’s earlier Veo model inside the Gemini app, marking a shift from standalone video generation toward what Google is calling “anything from anything” creation. What Gemini Omni actually does Early demonstrations showed effective text rendering within video, along with advanced scene editing capabilities. Google is emphasizing improvements in world understanding, physics simulation, and character consistency. The company drew comparisons to its Nano Banana image model, which earned praise for visual fidelity. Gemini Omni extends that same logic into motion and sound, wrapping everything into a conversational interface where users can iteratively edit and refine their clips through dialogue. Initial availability spans the Gemini app, Google Flow, YouTube Shorts, and additional tools for Google AI subscribers. The 10-second cap on clip length is expected to expand over time, though no specific timeline has been announced. From Veo to Omni: the lineage Google’s generative video efforts trace back to the original Veo model, which progressively gained features between 2025 and early 2026: native audio support, longer clip capabilities, and image-to-video functionality arrived in incremental updates. Veo was essentially a single-purpose tool. Omni represents a philosophical shift toward unified multimodal models, systems that reason across different types of media rather than treating each one