Gemini Omni Flash Is Live on Genfire: Generate, Reference, and Edit Video in One Place
Gemini Omni Flash is live in Genfire's Video Studio: text-to-video, image-to-video, multi-reference shots and chat-style video editing. How to use it.
It's Here. For Real This Time.
We covered the leak in May and the I/O announcement where we promised Omni would land in Genfire "the moment it's available to integrate."
That moment arrived. Gemini Omni Flash is now live in Genfire's Video Studio — text-to-video, image-to-video, multi-reference shots, and chat-style video editing, all routed through the same dispatcher that already handles Seedance 2.0, Veo, Kling, Sora 2, and 15+ other models.
No waitlist. No separate Google subscription. If you have Genfire credits, you can use Omni right now.
Update, September 2026: Gemini Omni Flash 1.1 has since shipped and is now the default Omni tier on Genfire. It adds a 360p / 720p / 1080p / 4K resolution ladder, an end frame on image-to-video, and reference video clips alongside reference images. Everything below still describes how the Omni family works; see the Gemini Omni Flash model page for the current specification.
What You Can Actually Do With It Today
Omni Flash on Genfire ships in four modes, each pickable from the Video Studio model card. Here's the real, shipped behavior — not the spec sheet.
1. Text-to-Video
Prompt in, video out. Pick 16:9 or 9:16 and a duration anywhere from 3 to 10 seconds (default 8). Pacing and audio are prompt-driven — there's no separate "generate audio" toggle, so you describe the soundscape in the prompt itself:
"...in a single continuous shot, include calm background music, no dialogue."
That's a feature, not a limitation. Omni was built to take direction in plain language, and Genfire passes your prompt straight through.
2. Image-to-Video
Drop in a single starting image plus a prompt, and Omni animates it. Same 3–10s duration range, same 16:9 / 9:16 aspect ratios. Great for turning a Nano Banana / GPT Image 2 still into a moving shot without leaving Genfire.
3. Reference-to-Video
This is the one most people sleep on. Feed Omni one or more reference images and point at each one with Google's zero-indexed tag — <IMAGE_REF_0>, <IMAGE_REF_1> — in the order you attached them, then say what each one is for: "The character in <IMAGE_REF_0> walks through the city in <IMAGE_REF_1>. Use Image1 as a character reference and Image2 as a background reference." (In Genfire's studio the @Image1 chips do this for you.) You get reference-guided shots where a specific character, object, or style is locked across the generation.
4. Edit Mode — the headline
Bring an existing clip, give Omni a one-line instruction, and it restyles or edits that clip — not a regeneration that hopes for the best. Edit mode takes only a prompt + your source video (no duration or aspect ratio — it inherits the source). Simple prompts work best, and the pro tip is to append "Keep everything else the same." to preserve the rest of the scene.
Heads up: Per Google's regional rules, edit mode (editing your own uploaded video) is not available in the EEA, Switzerland, or the UK, and voice editing is unsupported everywhere for now. Text-to-video, image-to-video, and reference modes work globally.
The Demo That Sold Everyone: Multi-Turn Editing
The single best argument for Omni is still the violin sequence — the same scene edited four times in a row, where every edit holds without the character or performance drifting. This is what "edit, don't regenerate" actually looks like:
Until now, every "edit" with Veo, Sora, or Kling meant regenerating from scratch and praying for consistency. Omni is genuinely editing. On Genfire, you chain these by running edit mode on the output of your previous generation — your gallery becomes the multi-turn thread.
What It Costs
Omni Flash bills per second of output, flat, with no resolution multiplier on the original endpoints:
| Mode | Credit Key |
|---|---|
| Text-to-Video | gemini_omni_flash |
| Image-to-Video | gemini_omni_flash_i2v |
| Reference-to-Video | gemini_omni_flash_ref |
| Edit | gemini_omni_flash_edit |
There is no premium tier surcharge and no per-resolution math on the original Omni Flash endpoints — Omni sits at the fast, low-cost end of Genfire's video lineup, exactly where Google slotted the "Flash" variant. The studio shows the exact credit total for your chosen length before you hit generate, and the pricing page carries the current rates.
Where It Sits Next to Seedance 2.0 and Veo
We've said it in both prior posts and it's still true: Omni doesn't try to beat Seedance 2.0 on raw photorealism. That's not its job.
| Gemini Omni Flash | Seedance 2.0 | Veo 3.1 | |
|---|---|---|---|
| Text-to-Video | ✅ | ✅ | ✅ |
| Image-to-Video | ✅ | ✅ | ✅ |
| Multi-Reference Composition | ✅ (described by position, in plain language) | Partial | Partial |
| Edit an Existing Clip | ✅ | ❌ | ❌ |
| World-Knowledge / Physics Grounding | ✅ (Gemini) | ❌ | Limited |
| Raw Cinematic Fidelity | Strong | Best in class | Strong |
| Speed / Cost | Fast, low-cost | Premium | Premium |
| Content Provenance | SynthID (default) | ❌ | SynthID |
The Genfire play is the same as it's always been: use each model for what it's best at. A typical workflow now looks like:
- 1Generate a cinematic establishing shot with Seedance 2.0 for top-tier fidelity.
- 2Send a clip to Omni Edit to change one thing — swap a background, restyle it "anime," remove an object — fast and cheap.
- 3Loop the result into a Storyboard, lip-sync, or ControlFoley sound pass.
No model lock-in. No re-uploading between tools. One subscription, every model.
How to Use It Right Now
- 1Open Video Studio in Genfire.
- 2Pick the Gemini Omni Flash card (it's badged NEW) and choose a sub-model: Flash (text/image-to-video), Reference, or Edit.
- 3For Edit, drop in your source clip; for Reference, attach your images and point at each one in the prompt with its
@Image1,@Image2chip (Genfire converts these to Omni's<IMAGE_REF_n>tags) — and say what each one is for. - 4Set aspect ratio (16:9 / 9:16) and duration (3–10s) where applicable, write your prompt, and generate.
A few prompting tips that match how Omni was trained:
- Direct the audio in the prompt — "include soft ambient music," "no dialogue," "footstep sounds in sync."
- Put negatives inline — "Do not show any text on screen."
- For edits, stay surgical — one change per instruction, plus "Keep everything else the same."
The Bottom Line
Two posts ago, Omni was a leak. One post ago, it was an I/O announcement we promised to ship. Today it's a live card in your Video Studio, and it does something no other model on the platform does — edit a clip you already have, by talking to it.
Omni joins the 20+ models already in Genfire, ready to slot into whatever your project needs. One subscription, every model, all the time.
Open Video Studio and run your first Omni clip — or create a free account if you're new. Starter credits included, no credit card required.
All video clips in this post are © Google / Google DeepMind, originally published with the Gemini Omni announcement on May 19, 2026. Embedded here for editorial commentary.