Nano Banana Pro Consistent Character: Thumbnail to Reel
Nano Banana Pro consistent character workflow: one face across your YouTube thumbnail and reel opening frame. Reference limits, prompts, and specs.
TL;DR
- A Nano Banana Pro consistent character setup starts with one locked reference set, not with a prompt. Everything downstream reuses that set.
- Nano Banana Pro accepts up to 14 reference images per generation, but only 5 of those slots carry character identity. Spend them on purpose.
- Generate the 16:9 thumbnail and the 9:16 opening frame from the same reference set, in parallel. Never crop one out of the other.
- Feed the 9:16 frame into an image-to-video model as the start frame, so the face that earned the click is the face that opens the video.
- Drift almost always shows up in three places: hair volume, jaw width, and wardrobe. I fix those in the prompt, not by adding more references.
The gap between your thumbnail and your first frame
Here is the failure I kept shipping. I would generate a thumbnail I liked, post it, then generate the video separately. The click-through was fine. The retention was not. Viewers landed on a first frame showing a person who was almost, but not quite, the person on the thumbnail. Slightly rounder face. Different hair. A jacket that changed color.
That mismatch reads as a broken promise before anyone hears a word of your voiceover. It is a continuity problem, and continuity is exactly what reference-image models are built to solve. The fix has been the same on every brand project I have delivered: stop treating the thumbnail and the video as two separate creative jobs. Treat them as two renders of one locked character.
What Nano Banana Pro actually gives you
Nano Banana Pro is Google’s name for the Gemini 3 Pro Image model. The number that matters for this workflow is the reference budget, and Google publishes it in the Gemini API image generation docs. You can mix up to 14 reference images into a single generation, but those 14 slots are not interchangeable.
| Reference slot type | Nano Banana Pro (Gemini 3 Pro Image) limit |
|---|---|
| Objects reproduced at high fidelity | Up to 6 |
| Characters held for identity consistency | Up to 5 |
| Style references | Up to 3 |
| Total mixed references per generation | Up to 14 |
Two other things from the same documentation are worth knowing before you build anything. Output runs at 1K, 2K, or 4K, which means you can render a thumbnail at a size that still holds up on a television screen. And every image the model produces carries a SynthID watermark, so treat these as clearly AI-generated assets and plan your disclosure accordingly.
The practical read on that table: character consistency is capped at five images. If you upload eleven photos of your character, you are not getting eleven slots of identity. You are getting five slots of identity plus a lot of noise competing for the object and style budget.
Step 1: Build the Nano Banana Pro character lock
I use five images and no more. The goal is not coverage of every possible angle. The goal is agreement. Five images that disagree with each other about hair length will produce a face that averages them badly.
My five slots:
- Front, neutral expression, flat lighting. This is the anchor. If only one image survives, it should be this one.
- Three-quarter left. Same session, same hair, same clothing.
- Three-quarter right. Mirror of the above.
- Close crop on the face only. Shoulders up. This carries skin texture and eye detail.
- Full body, standing, same wardrobe. This one anchors proportions so the model does not invent a different build when you ask for a wide shot.
All five come from the same shoot, the same day, the same lighting setup. If you are generating the character rather than photographing a real person, generate the anchor first, then use that anchor as a reference to produce the other four before you start the real work. Keep the files clean and reasonably sized, since oversized references do not buy accuracy.
Step 2: Generate the 16:9 thumbnail
Thumbnail first, because the thumbnail decides whether the video gets watched at all. I generate it at 16:9 natively rather than rendering something square and cropping.
YouTube’s own custom thumbnail documentation recommends 3840 x 2160 pixels with a minimum width of 640, at a 16:9 ratio, in JPG, PNG, or GIF. File size limits differ by upload path: up to 50 MB from desktop, 2 MB from mobile. For a Short, the same page points you to 9:16 at 2160 x 3840 instead.
Two prompt rules I follow for the thumbnail pass:
- Name the framing, not the emotion. “Chest-up, subject on the right third, looking directly at camera” beats “excited expression.” Emotion words push the model to reshape the face.
- Leave the text out of the render. Nano Banana Pro renders legible text well, but thumbnail copy changes more often than the image does. I keep the left third visually quiet and add the text in the editor so I can A/B the words without regenerating the face.
Step 3: Generate the 9:16 opening frame from the same lock
This is the step most people skip. They take the finished thumbnail and crop it vertically. That crop throws away the composition and usually cuts the subject badly.
Instead, run a second generation. Same five character references. Same style reference if you used one. New aspect ratio, new framing instruction. You are asking for the same person in the same world, staged for a vertical frame.
The prompt changes I make between the two passes:
| Element | 16:9 thumbnail pass | 9:16 opening frame pass |
|---|---|---|
| Framing | Chest-up, subject off-center | Waist-up or full body, subject centered |
| Negative space | Left third kept quiet for text | Top and bottom kept quiet for captions and UI |
| Eye line | Direct to camera | Direct to camera, or slightly off if the shot opens on movement |
| Background | Simplified, high separation from subject | Deeper, with something for the camera move to travel through |
| Character references | Same five | Same five |
Same wardrobe description in both prompts, written the same way. If the thumbnail prompt says “charcoal wool overcoat,” the vertical prompt says “charcoal wool overcoat.” Not “dark coat.” The model reads those as different garments.
Step 4: Animate the opening frame
Now the still becomes a video. I run the 9:16 frame through an image-to-video model as the start frame, which keeps the face I just locked instead of letting a text prompt reinvent it.
I do this inside Higgsfield because Nano Banana Pro and the video models sit in the same place, so the still moves into the animation step without a download-and-reupload round trip. Kling is my default for the motion pass when the shot has a person in it and the camera needs to behave. I wrote up the identity side of that separately in my Kling 3.0 character consistency workflow, and the two approaches stack: Nano Banana Pro locks the look, Kling holds it through the motion.
Runway still wins for me on one thing here, which is a clean, restrained camera move that does not fight the subject. When the shot is a slow push on a face and nothing else, that is where I send it. You can use code BYMICHYDEV25 on their pricing page.
Keep the motion small on an opening frame. One camera move, one subject action, nothing else. The first second of a reel is doing identity work, not choreography.
Where the Nano Banana Pro character still drifts
Even with a clean lock, three things slip. Here is what I change when they do.
| What drifts | What causes it | My fix |
|---|---|---|
| Hair volume and length | References shot on different days, or a style reference with strong hair of its own | Re-shoot the reference set in one session; drop the style reference for the character pass |
| Jaw and cheek width | Emotion adjectives in the prompt (“confident,” “intense”) | Replace emotion words with camera and framing language |
| Wardrobe color and material | Loose garment wording that changes between passes | Write one wardrobe string and paste it into every prompt verbatim |
| Age reading | Lighting shift between thumbnail and vertical pass | Describe the light the same way in both prompts, down to direction and softness |
If the face is still wrong after those, the problem is upstream. Go back and check whether your five references genuinely agree with each other. They usually do not.
A note on disclosure
Every image out of this pipeline is AI-generated and carries a SynthID watermark. YouTube asks creators to disclose realistic altered or synthetic content at upload, and the rules are laid out on its disclosure page. Instagram applies its own labeling. I am not your lawyer, and platform policy moves faster than any blog post. Read the current page for the platform you are posting on before you publish a face that could be mistaken for a real person.
FAQ
How many reference images should I actually use for a character?
Five, because that is the character-slot ceiling on Nano Banana Pro. Adding more images does not add more identity, and inconsistent extras make the result worse.
Can I use the thumbnail itself as the reference for the video frame?
You can, and it works better than nothing. It is still second-best. A generated image inherits the quirks of the generation that made it, so referencing it compounds those quirks. Reference the original lock in both passes instead.
Does this work for a real person rather than a generated character?
Yes, and it works better. Photographs of a real person, shot in one session, are the most internally consistent reference set you can build. That is the version I use for client work.
What resolution should I generate at?
Nano Banana Pro outputs 1K, 2K, or 4K. I generate the thumbnail at the highest tier available to me, because YouTube’s own recommendation is 3840 x 2160 and thumbnails now get viewed on television screens. The vertical frame does not need the same headroom, since the video model will re-encode it anyway.
Do I need a separate reel cover image?
Not with this workflow. The opening frame is doing the job, and it already matches the thumbnail because both came from the same lock. That is the whole point of running the two passes in parallel.
Where this fits in the bigger setup
This is one piece. If you want the b-roll side of Nano Banana Pro rather than the character side, I covered that in my Nano Banana Pro b-roll setup, which handles the shots with no person in them. And if you are applying this to client work, the vertical version of it shows up in my real estate video tour workflow.
Show Your Work
# SHOW YOUR WORK — Nano Banana Pro consistent character
REFERENCE LOCK (5 images, one session)
1 front neutral / 2 three-quarter left / 3 three-quarter right
4 face close crop / 5 full body standing
--- PASS 1: 16:9 THUMBNAIL ---
Chest-up portrait of the referenced subject, positioned on the
right third of the frame, looking directly at camera.
Wardrobe: charcoal wool overcoat, plain crew neck underneath.
Light: single soft key from camera left, gentle falloff, no rim.
Background: simplified dark interior, strong subject separation,
left third kept visually quiet. No text in image.
Aspect ratio 16:9. Highest available resolution.
--- PASS 2: 9:16 OPENING FRAME (same 5 references) ---
Waist-up shot of the referenced subject, centered in frame,
looking directly at camera.
Wardrobe: charcoal wool overcoat, plain crew neck underneath.
Light: single soft key from camera left, gentle falloff, no rim.
Background: same dark interior, extended depth behind subject,
top and bottom kept quiet for captions.
Aspect ratio 9:16.
--- PASS 3: ANIMATE ---
Start frame: output of Pass 2.
Motion: slow push in, 20 percent. Subject holds eye line,
one small breath. No other movement.
Tool: Higgsfield -> higgsfield.ai/?ref=michydevWhat to do next
Build the five-image lock before you write a single prompt. It takes one session and it removes the drift problem from every render that follows. Then run the two passes side by side and compare the faces at full size, not in a preview where everything looks fine.
If you want the full production system I run behind this, including the prompt structures, the shot templates, and the delivery workflow I use with clients, it is in the AI Video Blueprint — $27.
By Michele De Vivo — AI Video Producer. Last updated: August 7, 2026.
Want video like this for your business?
Tell me about your project. I usually reply in under 4 hours.
Get in touch