LiveGemini Omni 1.1 Flash is live: image references, 4K output and scene extensions up to 40 seconds. Try it now →

← All posts
Oct 4, 2026 · 11 min read

What Is Gemini Omni? Google's Any-to-Any AI Video Model Explained

Text, an image, an audio waveform and a film strip flowing into a single video frame, illustrating Gemini Omni's any-to-any model

Gemini Omni is Google DeepMind's family of "any-to-any" AI models: you give it any mix of text, images, audio and video, and it creates a finished video with sound. You can then edit that video by simply telling it what to change. The first model in the family, Gemini Omni Flash, launched at Google I/O on May 19, 2026. Its upgrade, Gemini Omni 1.1 Flash, has been generally available since August 27, 2026.

Google describes Omni as the point "where Gemini's ability to reason meets the ability to create." In practice, that means a video model that understands physics, history and context, not just pixels. It also means you can work with it the way you work with a chatbot: one request at a time, keeping what works.

This guide covers what Gemini Omni is, how it works, what each version can do, how it compares with Veo, what it costs, and where you can use it today, including on OmniRender, where Gemini Omni 1.1 Flash runs in your browser with a free plan.

Gemini Omni at a glance

Gemini Omni (1.1 Flash)
DeveloperGoogle DeepMind
AnnouncedMay 19, 2026 (Google I/O)
Current modelGemini Omni 1.1 Flash, model ID gemini-omni-1.1-flash, generally available since Aug 27, 2026
InputsText, images, video clips, audio (voice references in Google's apps)
OutputVideo with native audio: music, sound effects and lip-synced dialogue
Clip length3 to 10 seconds per generation; extend in 10-second steps up to 40 seconds
Resolution720p by default; 360p drafts; 1080p and 4K via upscaling
Aspect ratios16:9 (landscape) and 9:16 (vertical)
EditingMulti-turn conversational editing of generated or uploaded clips
ControlImage and video references, first and last frame, timestamps
WatermarkInvisible SynthID on every video, plus C2PA credentials in Google's apps
Where to use itGemini app, Google Flow, YouTube Shorts and Create, Google Vids, Gemini API, Google AI Studio, partner apps such as OmniRender

Sources: Google DeepMind model page, Gemini API docs.

How Gemini Omni works

Most AI video tools chain separate systems together: one model reads your prompt, another draws frames, a third adds sound. Gemini Omni does all of it in one model. That single design choice explains most of what makes it different.

Diagram: text, image, audio and video inputs feed into one model that outputs video with sound
Gemini Omni reads every input in one model and outputs video with native sound.

1. Native multimodality

Gemini Omni reads text, images, audio and video at the same time, in the same model. So when you write "make the character from this photo dance to the rhythm of this clip," Omni actually understands how the photo, the clip and the instruction relate. It doesn't just paste them together. The result is more consistent characters, more faithful references and audio that matches what happens on screen.

2. Gemini's world knowledge

Omni is built on the same foundation as Google's Gemini language models. It knows how gravity, momentum and liquids behave, and it knows a lot about history, science and culture. That's why a short prompt like "a claymation explainer of protein folding" can produce something accurate rather than just pretty. Google's prompt guide makes the point directly: older models like Veo needed precise instructions, while with Omni you can describe what you want and let the model fill in the details.

3. Conversation as the editing tool

Google compares Omni to Nano Banana, its popular image editor, applied to video. Each instruction builds on the last one. You can say "swap the red car for a vintage blue convertible," then "now make it rain," and Omni keeps the rest of the scene intact. Through the API this works with a previous_interaction_id, so the model remembers the video without you re-uploading it.

Three frames of the same coastal road scene: the original with a red car, the same scene in the rain, then the red car swapped for a blue convertible
Each edit builds on the last: the scene, camera and road stay the same while one detail changes. (Illustration)

4. Reasoning before rendering

Because Omni plans before it draws, it handles instructions that pure video generators struggle with: timed events ("after 3 seconds, a woman enters"), readable on-screen text, rapid-fire sequences and actions synced to music. By default it will even plan a few different shots to tell a short story. You can ask for one continuous shot if you prefer.

Key features of Gemini Omni

Text to video. Describe a scene, the camera move and the mood, and Omni returns a 3 to 10 second clip with sound. It understands film language: dolly-in, whip pan, locked-off shot, handheld.

Image to video. Upload a photo, product shot or even a sketch and Omni animates it. You can use the image as the literal first frame, or only as a guide for movement and style.

Reference to video. Combine several images (characters, outfits, products, a style) plus up to three short video clips of up to 3 seconds each. Omni keeps each reference consistent across the clip. The Gemini app accepts up to 5 photo references.

Native audio and dialogue. Music, ambient sound and speech are generated with the picture, not added afterwards. Put a line in quotes and a character says it with matching lip movement.

Conversational video editing. Change the lighting, background, outfit, camera angle or style of an existing clip with one sentence. Omni changes only what you ask and preserves the rest. You can also edit your own uploaded footage (with regional limits, see below).

Scene extension up to 40 seconds. New in 1.1 Flash: continue a clip from where it ends, 10 seconds at a time, up to 40 seconds in total. Omni reads up to the last 10 seconds of the clip to keep motion, characters and audio coherent.

First and last frame control. Supply a start image and an end image, describe the move between them, and Omni fills in the shot. It's ideal for transitions, orbits and seamless loops.

Draft in 360p, finish in 4K. Iterate cheaply at 360p, then render the keeper at 720p, or upscale it to 1080p or 4K.

Readable text and timing. Omni renders on-screen text accurately and follows timing cues like "[0-3s]" or "every 2 seconds, cut to a new location."

Real-world physics. Google trained Omni to handle gravity, momentum and fluid dynamics more convincingly, so chain reactions, splashes and fast sports read as real.

Gemini Omni versions: Flash vs 1.1 Flash (and Pro)

Gemini Omni is a model family, and so far Google has shipped two Flash models. "Flash" is Google's name for its fast, efficient tier.

Gemini Omni FlashGemini Omni 1.1 Flash
ReleasedMay 19, 2026 (apps); API preview June 30, 2026August 27, 2026 (generally available)
Model IDgemini-omni-flash-preview (retirement from Sept 30, 2026)gemini-omni-1.1-flash
Max length10 seconds per clip10 seconds per clip, extendable to 40 seconds
Resolution720p360p draft, 720p, 1080p and 4K (upscaled)
ReferencesImages, plus a source video to editImages plus up to 3 video reference clips (3 s each)
First/last frameNoYes
Scene extensionNoYes, 10 s steps reading up to 10 s of prior context

What about Gemini Omni Pro? Launch coverage pointed to a larger Omni model above Flash, but as of October 2026 Google has not released or dated one. We'll update this guide when it lands.

Gemini Omni release timeline

DateMilestone
Sept 30, 2026Preview model ID gemini-omni-flash-preview reaches its earliest deprecation date
Aug 27, 2026Gemini Omni 1.1 Flash released: scene extension to 40 s, first/last frame, 360p to 4K
June 30, 2026Developer access via the Gemini API opens in preview
June 2026Omni Flash reaches #1 in the Arena.ai Video Arena for text-to-video and image-to-video
May 19, 2026Gemini Omni announced at Google I/O; Omni Flash rolls out in the Gemini app, Google Flow and YouTube Shorts

Gemini Omni vs Veo: what's the difference?

Veo is Google's dedicated video generation model, built for cinematic shots from a prompt. Gemini Omni is a reasoning model that creates and edits video. Both come from Google DeepMind, and Veo still exists as a separate model for developers. But in the Gemini app, Omni has replaced Veo 3.1 as the default video model.

Gemini OmniVeo
What it isMultimodal reasoning model that outputs videoSpecialized text/image-to-video model
Best atEditing, combining references, story logic, timingSingle cinematic generations
Prompting styleDescribe intent; Omni infers details from world knowledgePrecise, detailed prompts work best
InputsText, images, video, audio in one requestText and images
Edits a finished clipYes, through conversation, turn after turnLimited
Status in the Gemini appDefault video modelReplaced by Omni

On quality, Omni Flash entered the Arena.ai Video Arena at #1 for text-to-video in June 2026, about 158 points ahead of Veo 3.1. Leaderboards move fast and newer rivals have since closed the gap, so treat rankings as a snapshot rather than a verdict.

The short version: choose Omni when you want to iterate, edit and mix inputs. Veo remains a solid option for developers who need a single, tightly specified cinematic shot.

Where to use Gemini Omni (and is it free?)

Gemini Omni is available in several Google products and through partner apps. Access and cost depend on where you use it.

WhereWho can use itCost
OmniRenderAnyone, in the browser, worldwideFree plan (a few 360p videos a month, no card); paid plans from $14/month
YouTube Shorts and YouTube CreateEveryoneFree
Gemini appGoogle AI Plus, Pro and Ultra subscribers (18+)Included in the subscription; features vary by tier and country
Google FlowGoogle AI Plus, Pro and Ultra subscribersIncluded, uses Flow credits
Gemini API / Google AI StudioDevelopers on the paid tierAbout $0.10 per second of 720p video
Gemini Enterprise Agent PlatformBusinesses on Google CloudEnterprise pricing

Is Gemini Omni free?

Yes, in a few places. YouTube Shorts and the YouTube Create app let anyone use Omni Flash at no cost, within YouTube's own tools. College students can also get Google's Pro plan free for a year. On OmniRender's free plan you get 160 credits a month, enough for up to 3 videos at 360p, with no card required.

The Gemini API has no free tier for Omni.

How much does Gemini Omni cost?

For developers, Google prices Omni by tokens: $1.50 per million input tokens and $17.50 per million video output tokens. One second of 720p video uses 5,792 tokens, which works out to roughly $0.10 per second, or about $1 for a 10-second clip (Gemini API pricing).

For creators, a subscription is usually simpler. OmniRender plans run from Creator ($14/month, up to 56 videos) to Studio ($49/month, up to 200 videos) and Agency ($99/month, up to 440 videos). Every paid plan includes a commercial licence, and one-time credit packs that never expire are available if you'd rather not subscribe.

Want to build Omni into your own product? See the Gemini Omni API guide.

Limitations to know before you start

Gemini Omni is impressive, but it has clear limits today. Knowing them saves credits.

  • Short clips. A single generation tops out at 10 seconds. Longer scenes come from chaining extensions, up to 40 seconds in total, and only at the end of a clip (no prepending or inserting in the middle).
  • Regional limits on uploads. In the EEA, Switzerland and the UK, you can't edit or extend videos you upload. Videos generated by Omni can still be edited and extended everywhere.
  • People and likeness. Uploading images of certain recognizable people isn't supported, and images of minors can't be uploaded or edited in the EEA, Switzerland and the UK.
  • Voice. Editing someone's voice isn't supported, and you can't add new dialogue when extending an uploaded clip where someone is already talking.
  • Video references. Omni uses video references for likeness only; their audio is ignored. Reasoning across multiple source videos at once isn't supported yet.
  • Language. English is fully supported. Other languages often work, but Google hasn't formally evaluated them.
  • No negative-prompt field. Write exclusions into the prompt itself, for example "no dialogue" or "no text on screen."

Safety and watermarking

Every video Omni creates carries SynthID, Google's invisible watermark. Content made in the Gemini app, Google Flow or YouTube also includes C2PA Content Credentials. You can upload a video to the Gemini app and ask whether it was made with Google AI. Before release, Google ran automated and human red-teaming (by specialist teams outside the model's development team) plus ethics and safety reviews.

What's next for Gemini Omni

Omni starts with video, but Google has said image and audio outputs are on the roadmap. Expect two directions in the coming months: a heavier model above Flash, and more output types from the same any-to-any system. For now, Gemini Omni 1.1 Flash is the production model, and the fastest way to learn it is to use it.

Gemini Omni FAQ

What is the difference between Gemini and Gemini Omni?

Gemini is Google's family of AI models and the name of its assistant app. Gemini Omni is a specific model family inside Gemini that creates and edits media, starting with video. You can use Omni inside the Gemini app, but it is a different model from the Gemini chat models.

How do I get Gemini Omni?

The fastest way is in your browser: create a free OmniRender account and Gemini Omni 1.1 Flash is the default model. You can also use it in the Gemini app and Google Flow with a Google AI Plus, Pro or Ultra plan, free in YouTube Shorts and YouTube Create, or through the Gemini API. Our step-by-step guide on how to use Gemini Omni Flash walks through each option.

Does Gemini Omni cost money?

It can be free: YouTube Shorts, YouTube Create and OmniRender's free plan all let you try it without paying. Heavier use needs a Google AI subscription, an OmniRender plan from $14/month, or paid Gemini API usage at about $0.10 per second of 720p video.

Is Gemini Omni available in my country?

Google is rolling Omni out wherever the Gemini app is available, with features varying by plan and region. OmniRender works worldwide in the browser. Users in the EEA, Switzerland and the UK can't edit or extend their own uploaded videos.

How long can Gemini Omni videos be?

Each generation is 3 to 10 seconds. With Gemini Omni 1.1 Flash you can extend a clip in 10-second steps up to 40 seconds in total.

Does Gemini Omni generate sound?

Yes. Music, sound effects and dialogue are generated together with the video. Put dialogue in quotes and the character speaks it with lip sync.

Can Gemini Omni generate images?

Not yet. Today Omni outputs video only. Google has said image and audio outputs will follow. For images, Google's Nano Banana models are the current option.

Is Gemini Omni open source?

No. Gemini Omni is a proprietary Google DeepMind model, available through Google's apps, its API and licensed partners.

What is Gemini Omni 1.1 Flash?

It's the current, generally available version of Omni (model ID gemini-omni-1.1-flash). Compared with the original Omni Flash it adds scene extension to 40 seconds, first and last frame control, video references and 360p-to-4K output. See our full Gemini Omni 1.1 Flash overview.

Try Gemini Omni now

The best way to understand Gemini Omni is to make something with it. Write one sentence, add a photo if you like, and see what comes back. Then tell it what to change.

Start creating free on OmniRender, then read the Gemini Omni prompt guide to get better results from every credit.

Sources

More like this

Reading about prompts is the slow way to learn prompts.

Try one right now