Gemini Omni 1.1 Flash is now available in ComfyUI via Partner Nodes. It ships in the Gemini Video Omni node, where you select “Omni Flash 1.1” from the model dropdown and get text-to-video, image-to-video, reference-to-video, video editing, and scene extension from a single node. Resolution runs from 360p, up to 4K, every clip comes with a generated audio track, and the node accepts an image and a video input alongside your prompt.
Model Introduction
Gemini Omni 1.1 Flash is Google’s multimodal video model, built for fast generation, editing, and cinematic control. It is an improvement from the earlier Omni Flash version.
It is natively multimodal: text, image, audio, and video are processed together, which is what keeps output cohesive across modes. It supports conversational editing: describe a change in plain language and the model applies it while leaving the rest of the clip alone. And it carries Gemini’s world knowledge, pairing an understanding of physics with context from history, science, and culture, so scenes read as plausible rather than just photoreal.
In ComfyUI, the node exposes a prompt, resolution (360p, 720p, 1080p, 4K), aspect ratio (16:9 or 9:16), a task type selector (text_to_video, image_to_video, reference_to_video, edit, or extend), and optional image and video inputs.
Model Highlights
Resolution up to 4K
Output is available at 720p, 1080p, and 4K. Iterate on prompt and composition at 720p, then rerun the keeper at 1080p or 4K for delivery.
Video editing
Omni 1.1 Flash edits video by instruction. Connect a clip to the video input, describe the change, and the model applies it while preserving everything else. Short prompts work best here: “make this video anime,” “change the lighting to be more dramatic,” “change the text on the sign to say ‘Omni Flash’.” Adding “keep everything else the same” helps hold visual consistency when you are targeting one element.
Edits can be stacked. Because the model retains context from the previous result, you can adjust lighting in one pass and swap the background in the next without re-describing the scene.
Scene extension
Extend an existing clip and the model reads the prior motion and composition to carry continuity, character identity, and lighting through the added footage. “Continue the shot: the woman finishes the violin solo and takes a bow” picks up where the previous clip ends.
Prompting Omni 1.1 Flash
The model does a lot of its work from the prompt, so a few behaviors are worth knowing before you start.
Scene structure. By default the model composes a short narrative with several shots. If you want one continuous take, say so explicitly: “in a single unbroken scene,” “in a single continuous shot,” or “no scene cuts.” Detailed prompts that cover the scene, camera move, lighting, and mood outperform short ones.
Image inputs. When you supply an image, the prompt decides how it is used. The model can animate a product shot or photograph directly, or treat a drawing purely as a motion guide and produce realistic footage that does not show the sketch at all. Pair high-resolution inputs with specific motion descriptions; “make it move” will not get you far. To pin an image to a role, use tags in the prompt: <FIRST_FRAME> makes it the opening frame, and <IMAGE_REF_N> (indexed from 0) makes it a reference for a style, character, or object.
Editing prompts. Short instructions work best; overly descriptive edit prompts tend to introduce unintended changes. “Make this video anime,” “change the lighting to be more dramatic,” or “change the text on the sign to say ‘ComfyUI’” is the right register. When you are targeting one element, add “keep everything else the same.”
Timing. Events can be placed in natural language (”after 3 seconds, a woman enters the scene,” “every 2s cut to a new frame”) or with timecode brackets like [0-3s], [3-6s].
Audio. Every clip gets a generated soundtrack. It responds to direction, especially for music: “include calm background music,” “the video has a high energy techno beat.”
On-screen text. Rendered text is readable and controllable, so if a sign, label, or caption will appear, spell out what it should say.
Negatives. There is no separate negative prompt field. Put exclusions in the prompt itself: “no dialogue,” “no extra sound effects.”
Getting Started
Update ComfyUI to the latest version, or access Comfy Cloud.
Find the Gemini Video Omni node via the Node Library, or load a template from the Templates panel.
Select “Omni 1.1 Flash” as the model, add your prompt and any image or video inputs, set resolution and task type, and run.
Try it on Comfy Cloud, and let us know what you build.
As always, enjoy creating!

