AI Video Is Now Iterated at 360p and Rendered at 4K: Gemini Omni 1.1 Flash's Draft-First Economics
media August 28, 2026 · Mintec

AI Video Is Now Iterated at 360p and Rendered at 4K: Gemini Omni 1.1 Flash's Draft-First Economics

Google shipped Gemini Omni 1.1 Flash on August 27, 2026: 10-second scenes that extend to 40 with continuity, first/last-frame control, video references, 4K upscaled output, and cheap 360p drafts. Here's the iteration math behind the draft-first workflow we run at Virtalio — and why it rewrites the economics of producing AI video for brands.

On August 27, 2026, Google shipped Gemini Omni 1.1 Flash — and the most important part of the announcement isn't 40-second video or 4K. It's that the cost of iterating AI video just split in two. There is now a draft resolution — 360p, cheap — and a delivery resolution — 4K, expensive. That split turns the correct production workflow into an arithmetic problem, and the arithmetic favors teams that iterate cheap and render expensive exactly once.

At Mintec we produce AI video daily for Virtalio and for agency clients. The question that changed this week isn't "does AI video look good yet?" — that was settled months ago. It's "how many iterations can you afford before the final render?" And that number just multiplied.

What actually changed in Omni 1.1 Flash

Google's official announcement frames 1.1 Flash around control and economics, not a magic photorealism leap. Five verifiable changes:

  1. 10-second scenes that extend to 40. The model generates clips up to 10 seconds and extends them in 10-second intervals, preserving the original clip's look, motion, and storytelling. The end result: up to 40 seconds of narrative flow.
  2. First and last frame control. You can fix how a scene starts and how it ends. This is what turns generation into direction — the model no longer unilaterally decides the opening or closing frame.
  3. Video references. Not just reference images: clips as anchors for style, light, and motion.
  4. 4K via upscaling. The detail most headlines missed: 4K output is upscaled, not natively generated — specialist coverage verifies it against the model's documentation. It matters for production — more on that below.
  5. Cheap 360p drafts. Google says it plainly: render fast, cheap 360p drafts before committing to the high-resolution render.

That fifth point is the quiet earthquake of the announcement, and almost nobody put it in a headline.

The cost curve split in two

Until this week, AI video was produced in a single mode: one prompt, one render, review the result, and if it doesn't work you pay for the next render at the same price. The cost of being wrong was constant and high. Iteration was expensive, so teams iterated little, accepted mediocre output, or went back to traditional production.

With two resolution tiers, the math flips:

WorkflowIterations per sceneCost per iterationHow you fix a mistakeCost of discarding a version
One-shot (straight to final render)1–2High, delivery resolutionRe-prompt and pay the full render again100% of the render
Draft-first (360p → 4K)10–20 drafts, 1 finalMinimal at 360pAmend the block, approve, upscaleCents, not dollars

The rule we apply at Mintec: final output quality correlates with the number of iterations you paid for, not the resolution of the first attempt. Twenty 360p drafts cost less than two discarded 4K renders, and the twentieth draft is almost always better than the first direct render. 4K doesn't fix a badly directed scene — it just makes it sharper.

How we produce at Virtalio: the workflow 1.1 Flash accelerates

At Virtalio we were already working with Gemini Omni under one principle: intention first, explicit 10-second segments, a reference image as anchor, and iterative amendments. Version 1.1 Flash doesn't force us to change the method — it makes every step of the method cheaper.

Our flow for a 30–40 second ad:

  1. Visual anchor. A reference image locks the character identity, palette, and style before a single motion prompt is written. Without an anchor, every generation is a new lottery; with one, every iteration is a version of the same piece.
  2. 10-second blocks. We don't prompt "a 40-second ad." We write four 10-second narrative blocks — hook, argument, demonstration, close — and extend them in sequence. 1.1 Flash's continuity-aware extension is exactly this mechanic, now with explicit context between blocks.
  3. Draft-tier iteration. The scene is amended at 360p until direction, framing, and pacing are right. On-screen text, fixed camera, natural movement: every correction costs a fraction of what it used to.
  4. Single final render. 4K — or the actual delivery resolution — is paid for once, on the already-approved version.

This flow is the difference between producing AI video and operating a slot machine. And it's why 1.1 Flash matters more for agencies and brands than for casual users: casuals generate once; teams iterate until approval.

Extending to 40 seconds: plan in blocks, not giant prompts

The jump from 10 to 40 seconds doesn't mean "write a longer prompt." The model extends in 10-second intervals with up to 10 seconds of prior context as continuity memory. That turns a long piece into a scriptwriting problem, not a prompting problem.

The method that works for us: each 10-second block must stand alone as a micro-scene — one subject, one action, one optional framing change — because the model preserves style and motion between blocks, but you still direct the content block by block. If block 2 fails, you don't discard the full 40 seconds: you re-generate block 2 and request the extension again.

This connects to something we've written before: a sales deck is not a video brief — a long piece needs a documented structure before you touch the tool. The 10-second block is the atomic unit of that structure.

What the announcement doesn't say

Three caveats we didn't see in the headlines, and they should temper production enthusiasm:

  1. 4K is upscaling, not native generation. For ads and social, irrelevant: the material is consumed at 1080p or below. For large screens or demanding post-production, test against the brief. Our rule: if the delivery format is social, the upscale layer is enough; if it's broadcast, validate frame by frame.
  2. The continuity window is finite. Extension context covers up to 10 previous seconds. In 40-second pieces, late blocks depend on early blocks being well directed — another argument for the visual anchor and block-level scripting.
  3. Draft-first doesn't eliminate model selection. There are still one-take models, fine-control models, and models with dubious claims. Before committing pipeline architecture, we run the verification protocol we published with the Wan 3.0 case — and the exit plan we covered when Sora shut down its API remains mandatory with any vendor.

When draft-first, when one-shot

The new economics don't kill the old mode — they reclassify it:

  • Draft-first wins when client approval, brand consistency, or a complex brief is involved: cheap iteration absorbs the cost of being wrong, and 4K is reserved for the approved version. It's our default at Virtalio.
  • One-shot still wins when the format is short and disposable — test clips, creative-testing volume — or when the piece needs narrative and audio continuity in a single breath. For that call we keep the framework from Seedance 2.5 and single-take generation: one take isn't better or worse, it's the right tool for certain scripts.

The short answer: if the piece passes through human eyes before publishing, iterate at 360p. If it goes straight to test without review, generate one-shot and measure.

Cost is no longer the excuse

For two years, the standard objection to AI video in production was the cost of iterating: "every change costs another render." By separating draft from delivery, Omni 1.1 Flash removes that objection from the conversation. What's left exposed — and this is what separates teams — is whether you have a direction method, block-level scripting, and an anchor image, or whether you're still re-prompting blind.

Iterating cheap doesn't save you from directing badly. It just charges you less for learning.

Frequently Asked Questions

What is Gemini Omni 1.1 Flash?

It's Google's video generation and editing model, announced August 27, 2026. It generates scenes up to 10 seconds, extends them in 10-second blocks up to 40 seconds while preserving visual and narrative continuity, supports first/last-frame control, video references, 4K output via upscaling, and cheap 360p draft renders.

Why render 360p drafts first?

Because final AI video quality tracks the number of iterations you can afford, and per-iteration cost collapses at draft resolution. Iterating at 360p costs a fraction of a 4K render: you can afford 10-20 attempts at a scene for the price of a single final-quality render, and reserve 4K for the version that's already approved.

How does scene extension to 40 seconds work?

You generate a base scene of up to 10 seconds, then request 10-second extensions, each informed by the previous context (up to 10 seconds of continuity). The model preserves the original clip's look, motion, and storytelling, so you plan a 40-second piece as a sequence of coherent blocks instead of one giant prompt.

Related Articles