Google Gemini 1.1 Flash: 4K Upscaling & Advanced Scene Control

Google Releases Gemini Omni 1.1 Flash for Directable Native Multimodal Video Generation

Google has released Gemini Omni 1.1 Flash (gemini-omni-1.1-flash), a production update to its native multimodal video generation and editing model. According to official release details, the update shifts Omni from a basic video generator into a directable tool that allows scene extension, precise frame pinning, and character consistency via video references.

The model is built on three core properties distinguished by Google from prior video architectures: native multimodality processing text, image, audio, and video together; conversational editing via the Interactions API; and world knowledge inherited from the broader Gemini ecosystem. Editing is stateful, meaning users pass a previous_interaction_id so the model applies updates without requiring a re-upload of the prior video.

Scene Extension and Context Windows

The primary upgrade in Omni 1.1 is its ability to analyze up to 10 seconds of prior context when continuing a clip, a significant jump from previous models that referenced only the final second, according to Google. Scene extensions run in 10-second increments up to a cumulative 40 seconds, with the model generating a 3 to 10 second continuation per API call. The system automatically edits some final frames of the input to ensure a continuous seam.

However, specific constraints apply to the extension pipeline. Extensions append exclusively to the end of a clip—prepending and mid-clip insertions are not supported. Uploaded input videos must measure 10 seconds or shorter unless users are extending a model-generated video across multiple turns. Furthermore, users cannot add new dialogue when extending an uploaded video featuring a speaker, though spoken dialogue is supported during multi-turn extensions via the Interactions API.

Pro Tip: When using the draft-then-upscale loop, set your resolution parameter to 360p for initial generation. Google reports that 360p previews run up to 60% faster and cost one-third of 720p processing, allowing you to iterate cheaply before rendering your final output in 1080p or 4K.

Keyframe Control and Video References

Omni 1.1 introduces strict shot-level camera and character control through first and last frame interpolation, enabling techniques like orbits, dolly-zooms, and seamless loops. Prompts bind media to specific roles using tags such as <first_frame> and <last_frame>.

For character consistency, video references accept a maximum of three clips lasting up to three seconds each. Google notes that audio inside a video reference is ignored, and reasoning across multiple videos is unsupported as it may degrade output quality.

Pricing, Cost Control, and Production Availability

Input pricing is set at $1.50 per 1 million tokens across text, images, video, and audio. Output pricing runs $9.00 per 1M text tokens and $17.50 per 1M video tokens. According to system specifications, video billing operates at 5,792 tokens per second of 720p output, translating to an effective cost of roughly $0.10 per second.

Every generated video includes an invisible SynthID watermark for programmatic provenance detection. Noted technical gaps in the current release include a lack of system instructions, temperature controls, top_p parameters, stop sequences, and native negative prompts—meaning negative constraints must be entered directly into the prompt text. Additionally, voice editing, audio references, and YouTube URLs as sources are unsupported.

Google has confirmed that Adobe (Firefly), Figma Weave, GMI Cloud, and Runway are already utilizing Omni Flash in production environments. The model is also available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, while living natively inside Google Flow for AI Plus, Pro, and Ultra subscribers.

Frequently Asked Questions

What is the maximum scene extension length for Gemini Omni 1.1 Flash?

Scene extensions run in 10-second increments up to a cumulative 40 seconds total.

Google Gemini Omni 1.1 Flash | The Biggest AI Video Update Yet? #aivideo

How does the cost savings work with 360p drafts?

According to Google, 360p previews generate up to 60% faster and at one-third the cost of 720p generation, making the draft-then-upscale approach the recommended production pattern.

Are free tiers or provisioned throughput available?

No. Omni 1.1 Flash is available strictly on paid tiers through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

Can I use YouTube URLs as a video source?

No. YouTube URLs cannot be used as a source input for the model.


Looking to promote your product release, GitHub repository, or technical webinar? Connect with our team to explore partnership opportunities.

Leave a Comment