Google’s Gemini “Omni” Leak Signals Video Is Moving Into the Assistant Layer
Google’s Gemini “Omni” Leak Signals Video Is Moving Into the Assistant Layer
May 6, 2026
Google appears to be testing built in video generation inside Gemini, with leaked interface screenshots showing a new video tab labeled “Powered by Omni.” If that holds, this is more than one more shiny generative media update. It suggests Google wants video creation to live inside the same assistant workflow already used for writing, research, image generation, and task orchestration. The leak was first surfaced in reporting from TestingCatalog, and it points to a bigger shift: AI video may be moving from specialty tool to embedded workflow layer.
That distinction matters. A standalone video model can be impressive. A video model inside Gemini changes the daily operating system for marketers, creators, and teams trying to move from brief to assets without opening six tabs and losing the plot in the process.
The headline is not just “Google might launch another video model.” The headline is that video generation may be getting folded directly into the assistant interface where creative work already starts.
What the leak actually shows
The current evidence is still leak territory, not official product documentation. Screenshots circulating online show Gemini’s “Create Videos” area with the phrase “Start with an idea or try a template. Powered by Omni.” That wording strongly suggests an internal model or system name, and it appears to point to a new video generation surface inside Gemini.
What we can reasonably infer is simple:
- Google is actively testing a native video creation surface inside Gemini
- “Omni” is likely the working name for the model or orchestration layer behind it
- The feature is being designed as part of Gemini’s core experience, not as a confirmed handoff to a separate public app
What we cannot confirm yet is whether Omni is a brand new model, a rebrand or wrapper around existing Veo related capability, or a broader multimodal system that could eventually connect text, image, and video generation under one roof. Right now, those remain plausible interpretations, not confirmed product facts.
Why this is a bigger deal
Google already has generative media assets across Gemini, Veo, Imagen, Flow, and its developer stack. The problem has never been “can Google make AI media.” The problem has been product sprawl. Great demos, slightly messier workflow reality.
If Omni brings video directly into Gemini, Google may be trying to solve the actual bottleneck: handoff friction.
For non technical teams, handoff friction is the tax you pay every time work moves from one tool to another:
- brief in one app
- image prompts in another
- video generation elsewhere
- feedback buried in chat
- final files lost in a folder called FINAL_final_v9_forreal
Putting video inside Gemini reduces that mess. In theory, a team could ideate a campaign, draft scripts, generate stills, and produce early motion concepts in one conversational flow. That does not magically make the outputs production perfect. But it does make iteration faster, which is where a lot of creative advantage lives now.
What this could unlock
If the leaked interface becomes a public feature, the real value is not just text to video. It is context aware video generation inside an assistant that already knows the brief, the audience, the messaging, and potentially the supporting assets.
Less context switching
Marketing teams could move from “give me three ad concepts” to “turn concept two into a six second teaser” without re briefing the system from scratch. That sounds small until you realize how much time creative ops loses to repetition.
Faster revision loops
Inside a chat based workflow, users could refine outputs conversationally:
- make it brighter
- shorten the pacing
- swap the product angle
- make this feel less corporate and more human
Yes, that last instruction will definitely be abused. But the workflow logic is sound.
Cross format continuity
If Gemini can carry context from copy to image to video, brands get a shot at more consistent campaign assets across formats. Not perfect consistency. But better than the current “every model interprets the brief like it skimmed the group project notes” problem.
| Potential capability | What it changes | Why teams care |
|---|---|---|
| Video inside Gemini | Fewer tool handoffs | Faster concept to draft workflow |
| Shared prompt context | Better alignment across assets | Cleaner campaign consistency |
| Template based creation | Easier repeatability | More useful for teams, not just tinkerers |
Can it be automated?
This is the adult question, and right now the answer is: not yet confirmed for Omni specifically.
Google does have an existing developer path through the Gemini API, and Google publicly documents video generation through Veo models in the Gemini API and Vertex AI. Public docs reference Veo 3.1 preview models for text to video and image to video generation, rather than any public “Omni” endpoint. That matters because Google clearly has the platform plumbing to expose new capabilities programmatically when it wants to. But there is currently no public documentation for an “Omni” video API endpoint.
So here is the practical read:
| Question | Best current answer | Meaning |
|---|---|---|
| Can you use it manually? | Possibly, if it launches in Gemini | Good for ideation and draft creation |
| Can you automate it? | Unknown for Omni; Veo video generation does have a public API path | Needs a public Omni API or equivalent developer surface |
| Can it plug into your stack? | Not proven yet | Workflow value depends on exposure beyond the app |
If Omni ships only as an in app feature, it will still be useful. But usefulness and automation are not the same thing. A closed UI makes for a fun demo. An API makes for infrastructure.
What looks real versus hype
There is enough in the leak to take the story seriously. There is not enough to assume Google has solved video production overnight. Let’s keep the grown up supervision on.
What looks real
- Native placement in Gemini: the UI leak suggests this is being folded into the assistant, not bolted on from the outside
- Template led workflow: “start with an idea or try a template” hints at practical user onboarding, not just raw prompting
- Multimodal direction: the Omni name suggests Google is pushing toward more unified media workflows, though the exact scope is still unconfirmed
What still needs proof
- Output quality: no confirmed Omni specs yet on duration, resolution, fidelity, or motion consistency
- Editing depth: we do not know whether this is simple prompt to clip or something closer to iterative post production
- Audio support: no public evidence yet that Omni itself supports generated audio, even though Google’s Veo 3 family does have public audio capable variants
- API availability: still the biggest question for teams building repeatable systems
- Production readiness: no public evidence yet on quotas, governance, enterprise controls, or rollout scope
Translation: this is promising workflow news, not a reason to fire your editor, your producer, or your common sense.
Why marketers should care now
Because even if Omni launches in a limited form, it points to where creative tooling is going: toward assistant native asset generation. The future is less “open a separate AI tool for every media type” and more “work inside one intelligent layer that can draft across formats.”
For executives, that means the strategic question is shifting from Which single model is best? to Which platform best reduces workflow drag while staying open enough to automate?
For marketers and creators, the opportunity is even more immediate. If video generation becomes a natural extension of the campaign planning process, teams can test more concepts, generate more first drafts, and move human attention upstream toward strategy, judgment, and taste. That is exactly where people should be spending their time anyway.
If you want broader context on how Google’s stack has been moving toward more workflow centric multimodal AI, our earlier COEY coverage of Gemini 3.1 Ultra and Veo 3.1 Lite tracks the same bigger pattern.
What to watch next
The next signal that matters is not more leak screenshots. It is official product detail.
Teams should watch for four things:
- Public announcement language: is Omni a model, a feature, or a wrapper around existing video systems?
- Access path: Gemini app only, Labs preview, Workspace rollout, or developer API exposure
- Control surface: templates, editing, reference images, audio support, aspect ratio options
- Workflow hooks: export options, structured metadata, enterprise controls, and eventually API access
Bottom line: the Gemini Omni leak matters because it suggests Google is pulling video generation into the assistant layer, where creative work increasingly begins. That is strategically more interesting than another standalone model launch. If Google follows through with usable controls and a real developer path, Omni could become a meaningful collaborator in content workflows. If it stays locked inside a glossy app tab, it will still be helpful, but mostly as speed, not system. And in 2026, the winners are not the teams with the most AI tabs open. They are the teams that turn AI capabilities into repeatable creative operations.





