Google Ships Two Lightweight Creative Models: Images in 4 Seconds, Video in 10
Google releases the Nano Banana 2 Lite image model and the Gemini Omni Flash video model, built for speed and savings — $0.034 per image and $0.10 per second of video, already embedded in the Gemini App and YouTube Shorts for free.


Google Ships Two Lightweight Creative Models: Images in 4 Seconds, Video in 10
Google releases the Nano Banana 2 Lite image model and the Gemini Omni Flash video model, built for speed and savings — $0.034 per image and $0.10 per second of video, already embedded in the Gemini App and YouTube Shorts for free.
On June 30, Google launched two lightweight creative models: the image model Nano Banana 2 Lite (technical codename Gemini 3.1 Flash-Lite Image) and the video model Gemini Omni Flash. One is built for speed, the other for savings, and both are already embedded in the Gemini App, Google Flow, YouTube Shorts, and YouTube Create, free to use. If you build AI applications or run an image-and-video production workflow, the core problems these two models solve are "is it fast enough, is it cheap enough, and can it slot into the pipeline I already have."
Model 1: Nano Banana 2 Lite — a 1K Image in 4 Seconds
This is the newest spin-off of the Nano Banana family and the fastest, cheapest model in the lineup. A 1K image in 4 seconds, at a unit price down to $0.034. It's positioned for rapid visual sketching and high-frequency A/B testing — if your AI app makes users wait 15 seconds per click for an image and the churn hurts, this model exists to fix exactly that.
On performance it scored an Elo of 1255, ranking fifth, ahead of Nano Banana Pro. It fits rapid visual sketching and high-frequency image testing.
Model 2: Gemini Omni Flash — Conversational Video Editing
The first member of the brand-new "Omni" family, positioned as conversational video editing. You can make partial modifications while preserving the original footage's camera work and pacing. Through the Interactions API, up to three consecutive editing rounds are supported, each retaining the context of the previous ones.
In Google's official demo, a woman films herself with her phone while Omni Flash layers effects onto the video step by step — pulling 3D balloon lettering out of the phone screen, pouring water from the screen "into" a glass — three editing rounds, each built on the last, with characters, lighting, and physics kept consistent.
The differentiator is the word "unified": text, images, audio, and video are processed as one multimodal understanding stack, backed by Gemini's built-in world knowledge and distribution reach straight into billion-user traffic entrances like YouTube Shorts and Search.
Two Models, One Pipeline
Google designed the two models to chain together:
- First use Nano Banana 2 Lite to produce a character key image in seconds
- Then feed it to Omni Flash to "bring the still to life"
- Then revise across multiple rounds in natural language
Google built three demo apps to show off the workflow:
- Anywhere (swap selfie backdrops with landmarks + generate video): upload a selfie, NB2 Lite drops you in front of landmarks worldwide, and one tap on Omni Flash turns the still into moving video
- Space Lift (room design proposals + video previews): snap a photo of your room, get multiple interior design schemes generated automatically, then convert the one you pick into a video preview with one click
- Omni Product Studio (product images to e-commerce video): turn static product shots directly into e-commerce-grade video ads
Price Comparison
On the image side, the Nano Banana family has three clearly differentiated tiers:
| Model | Unit price | Speed | Best for |
|---|---|---|---|
| Nano Banana 2 Lite | $0.034/image | 4 seconds | Quick sketches, A/B testing |
| Nano Banana 2 | $0.067/image | 4-8 seconds | Balance of quality and editing |
| Nano Banana Pro | $0.134/image | 10-20 seconds | Complex scenes, maximum detail |
Against OpenAI's GPT Image 2 (token-priced, roughly $0.053 per 1024x1024 image at medium quality and about $0.211 at high quality), NB2 Lite wins on both price and speed.
On video, Omni Flash is priced at $0.10 per second, going up against Veo 3.1 Fast (same price) and OpenAI Sora 2 Standard's 720p tier (also $0.10 per second). In a broader comparison, 5-second videos from ByteDance's Jimeng and Kuaishou's Kling generally run about $0.40, which works out to roughly $0.1 per second — so Omni Flash's "value" badge is nothing special on this dimension.
Known Weaknesses
Omni Flash's own documentation admits several limitations:
- No audio reference support
- No scene extension
- Video references under 3 seconds are "accepted by the API but not yet handled well by the model"
Community feedback adds more: rendering Chinese text in complex layouts still produces errors, generated images occasionally still grow a sixth finger, queues during peak hours stretch past 30 seconds — at odds with the "image in 4 seconds" banner — and style transfer for watercolor and oil painting remains shaky. Adobe has already plugged both models into Firefly, with WPP, Figma, Invideo, and Artlist as early adopters.
Who Should Use Them
- AI application developers: anyone who cares about image latency and unit price and needs the models embedded in a product workflow
- Content creators: rapid sketching, turning stills into video, iterating on visual concepts
- E-commerce/marketing teams: product images to e-commerce videos — the Omni Product Studio demo workflow can be adopted as is
Flagship models solve the capability-ceiling problem; these two lightweight models solve a different one — "is it fast enough, is it cheap enough, can it fit the workflow I already have." In the AI business, the long-run winner is usually not the one that runs highest, but the one that runs into your workflow first.
Related articles

Meshy: Turning Text and Images into Usable 3D Models in 20 Seconds
AI text/image-to-3D tool with a two-step workflow and FBX/OBJ/GLB/STL export, plus an official API and platform plugins.

Qwen-Image-3.0: Alibaba Qwen's Third-Generation Image Generation Model
Up to 4.5k tokens of input, crisp rendering of 10px-scale small text, and 12 languages — renders complex layouts like posters, exam papers, and storyboard comics in a single pass.

Build an Automated Reddit Overseas Customer Acquisition Workflow with Qwen 3.8
An eight-step fully automated pipeline: Apify scraping plus Qwen 3.8 scoring/drafting/self-review/retrospective — 260 posts filtered down to 22 opportunities, at a tenth of the cost.

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Qwen3.8-Max Preview: A Hands-On Early Access Guide
Qwen's 2.4T-parameter flagship preview is live on Token Plan, Qoder, and the Qwen website; officially rated second only to Fable 5 overall.

Making a Full-Scene Infographic with SenseNova U1 Pro
A hands-on tutorial for SenseTime's flagship multimodal model U1 Pro: turn raw data into a deliverable 8K infographic and full-match panoramic visual, automatically.