Google DeepMind released Gemini Omni 1.1 Flash on August 27, 2026, and the update brings substantial improvements to the platform's generative video model. The original Gemini Omni was promising but limited: clips were short, context awareness was minimal, and output resolution capped at 720p. Version 1.1 Flash addresses all three issues and adds new creative controls. If you've been waiting for Google's AI video tool to mature, this update is the one to look at.
Five core changes define the release: longer scene extensions with deeper context, first and last frame interpolation, 4K upscaling, a faster 360p preview mode, and video references in multimodal input. Let's walk through what each feature does and why it matters.
Scene Extension Gets a Major Upgrade in Gemini Omni
The previous version of Gemini Omni could analyze 1 second of prior video context when extending a clip. That's barely enough to understand what's happening in a scene. Version 1.1 Flash pushes that to 10 seconds, which means the model has real context about the action, lighting, and composition it should continue.
You can now extend videos in 10-second increments, pushing total length up to 40 seconds. For comparison, the old limit effectively kept you at around 10-second clips. That difference matters. A 40-second clip with 10 seconds of context is long enough for short-form social content, product demos, or storyboard prototyping. You're no longer stuck with fragments that need heavy editing to feel coherent.
For developers building on the Gemini API, this changes the scope of what's possible in a single generation session. Scene extension was the most requested feature according to the product team, and the jump from 1 second to 10 seconds of context is the kind of change that turns a demo tool into something you can actually ship with.
First and Last Frame Interpolation for Smooth Transitions
This feature lets you specify a starting frame and an ending frame. The model then generates the motion between them. If you've ever wanted to create a smooth visual bridge between two distinct shots without manual keyframe animation, this solves that problem directly.
The use case is specific but valuable. You might have a product shot at the beginning of a video and a different angle at the end. Instead of a hard cut or a crossfade, the model generates intermediate frames that connect the two states. It's a time-saver for anyone who'd otherwise be doing this work by hand in After Effects or DaVinci Resolve.
Product managers Anish Nangia and Alisa Fortin noted that this feature came directly from creator feedback. People wanted control over how scenes start and end, not just what happens in the middle. That feedback makes sense when you think about how much of video editing is about transitions rather than content.
4K Upscaling for Higher Quality Google AI Video
Previous outputs from Gemini Omni capped at 720p. That resolution is fine for prototyping and internal review. It doesn't work for final delivery on most platforms where viewers expect at least 1080p. The 1.1 Flash update adds 4K upscaling, so you can generate polished 1080p or 4K outputs from your clips.
This matters most for creators who publish directly to YouTube, Instagram, or other platforms where resolution affects how content is perceived. A 4K output from a generative video model won't match footage from a professional camera. But it's a significant jump from 720p for most viewing contexts, and it means you don't need a separate upscaling step in your pipeline.
The upscaling applies to all clip types in the workflow. Content created through scene extension and frame interpolation can all be rendered at 4K. You generate at standard resolution first, then upscale as a final step.
360p Preview Mode Cuts Cost and Iteration Time
When you're iterating on a prompt or testing different scene directions, full resolution is a waste. The new 360p preview mode generates previews up to 60% faster than 720p and costs roughly one-third as much. That's a big deal for anyone paying for Gemini API usage.
If you're running 50 iterations to nail a scene, the cost difference between 360p and 720p adds up fast. The speed improvement also means you can try more variations in the same window of time. For prompt engineering, where the difference between a good output and a great one often comes down to trying dozens of slight variations, faster and cheaper previews are the most practical improvement in this release.
The workflow is straightforward. Iterate in 360p to find your direction. Once you've found the version that works, switch to 720p or 4K for the final render. The two modes are separate API calls, so you're never locked into one resolution.
Video References in Multimodal Input for Gemini Omni
Gemini Omni 1.1 Flash now accepts up to 3 seconds of video as part of your multimodal input. You can reference an existing clip when prompting the model to generate new content. This is different from scene extension, which continues an existing video. Video references let you guide the style and motion of new generations.
If you've ever struggled to get a generative video model to match a specific visual style through text alone, video references solve part of that problem. Text prompts are still the primary input. But a 3-second video reference gives the model concrete visual information to work with, which helps when you're trying to match a particular aesthetic or motion pattern.
Here's how the key specs compare between the two versions:
| Capability | Previous Version | Gemini Omni 1.1 Flash |
|---|---|---|
| Prior context for extension | 1 second | 10 seconds |
| Maximum video length | ~10 seconds | 40 seconds |
| Preview modes | 720p only | 360p + 720p |
| Video references in input | Not supported | Up to 3 seconds |
| Maximum output resolution | 720p | 4K |
| Frame interpolation | Not available | First and last frame |
What This Update Means for Developers and Creators
The 1.1 Flash update moves Gemini Omni from an interesting experiment to a practical production tool. Developers building on the Gemini API get longer outputs, lower iteration costs, and new creative controls. Content creators get access through consumer-friendly interfaces without writing code.
The developer who was hitting the old limits benefits the most. Going from 1 second of context to 10 seconds, and from roughly 10-second clips to 40-second clips, changes what kind of content you can produce in a single session. The 360p preview mode makes experimentation affordable enough to use regularly. And video references add a creative direction tool that text-only prompting can't match.
This is still a generative video model. Output quality depends on your prompts and how well you use the new controls. If you're expecting to type a sentence and get a finished video, you'll be disappointed. But if you're willing to iterate and use the preview mode to test variations, the tools are meaningfully better than the original release.
Mobile Access Through the Gemini App on Android
Scene extension is available in the Gemini app for subscribers on Android. You can try the feature on your phone without writing a line of code. The full set of capabilities, including 4K upscaling, video references, and the Agent Platform API, is available through Google AI Studio for developers.
Google Flow, the dedicated creative interface for Gemini Omni, is available for AI Plus, Pro, and Ultra subscribers globally. That covers the web experience. On mobile, the Gemini app handles consumer-facing features. Scene extension in the app lets you extend clips directly from your phone, which is useful for quick iterations when you're away from your desk.
If you want to try the new AI video generation features on your phone, download the Gemini app from APKPure to get the latest version. The app receives regular updates with new Gemini capabilities as they roll out. As with any AI tool, review the privacy permissions and data usage settings before signing in with your Google account.



