Google Debuts Gemini Omni: AI That Creates Video from Inputs
Google has officially introduced Gemini Omni, a new suite of multimodal models designed to process and synthesize various inputs—including text, audio, images, and video—into cohesive, high-quality video content. Unveiled at the company’s I/O developer conference, the technology marks a significant evolution in Google’s goal to create a single neural network capable of reasoning across different media formats.
Unlike previous tools that simply stitched different media together, Gemini Omni is engineered to understand the nuances of physics, science, and history to produce consistent outputs. Sundar Pichai, CEO of Google, described the development as a shift from traditional text prediction to a form of reality simulation.
Key Features and Capabilities
The Omni family is designed to handle complex creative tasks through natural language prompts. Key capabilities include:
- Multimodal Reasoning: Users can input combinations of text, audio, and images, which the model interprets to generate accurate video sequences.
- Direct Editing: Images can be modified using simple text commands, removing the need for traditional, complex editing software.
- Digital Avatars: The platform allows for the creation of personalized digital avatars. To mitigate risks like deepfakes, users must complete an onboarding process that includes recording their voice and appearance.
- Watermarking: All content generated via Omni will feature Google’s SynthID digital watermark to verify its AI-generated origin.
Nicole Brichtova, director of product management at Google DeepMind, emphasized that this release represents the integration of Gemini’s intelligence with advanced media rendering capabilities. For instance, a simple prompt such as “a claymation explainer of protein folding” can result in a fully rendered video with synthesized voice-over, as demonstrated by DeepMind’s chief technologist, Koray Kavukcuoglu.
Rollout: Gemini Omni Flash
The first model in this series, Gemini Omni Flash, is available starting today. It is being integrated into the Gemini app, YouTube Shorts, and the AI creative studio, Flow. Currently, the model is capable of rendering 10-second video clips. Brichtova noted that this duration is a strategic choice for initial consumer adoption, with plans to support longer videos in the near future.

While the immediate focus is on consumer-friendly applications—such as creating “personalized memes” or editing vacation footage—Google intends to expand the utility of the model significantly. An API for Gemini Omni is scheduled for release in the coming weeks, targeting filmmakers, advertisers, and content creators who require more robust production workflows.
Looking ahead, Google plans to launch the Omni Pro model, which is expected to offer enhanced performance for professional-grade tasks. According to Brichtova, the Pro version will be deployed once it provides a distinct “step change” in capabilities over the Flash version.
The broader vision for the Omni architecture includes future iterations capable of generating images from audio or audio from video, further blurring the lines between different types of digital content creation.