How I Made a 1.4M-View AI Reel with ElevenLabs + MidJourney (Step-by-Step)
3-hour breakdown of the 1.4M-view reel. Every prompt, every visual beat, every edit decision. $3.20 in actual costs. See the full breakdown inside.
- 3 hours total work — $3.20 in tool credits. That’s it.
- 30 minutes was the question discovery (4-prompt chain), the rest was execution
- 168-word script, 47 seconds final, single idea per video — no filler
- 6 visual beats: mix of MidJourney + Higgsfield Soul + Kling motion
- Posted same export to YouTube/Instagram/TikTok simultaneously
The 1.4M-view reel on my channel took 3 hours to make. Here’s the exact step-by-step, with the actual prompts, the actual frames, and the decisions I made along the way.
It was a 47-second reel about cinematic AI video creation. The hook line: “I made 1.4 million views with zero cameras.” The visual concept: an empty film set at dawn with no equipment visible — implying that “cinema” no longer requires cameras.
Why those numbers matter: the algorithm gave it sustained reach because viewers kept WATCHING through second 47. Retention is the metric.
Step 1 — Finding the question (30 minutes)
I ran the 4-prompt question discovery chain (full framework here, free). Each prompt narrowed the field until I had a question worth answering.
Prompt 1 — The question I picked
“How are people making cinematic AI videos that don’t look cheap or uncanny?”
Source: Reddit r/aivideo, asked weekly across 6+ months. Top Google result was an 8-month-old tutorial using since-deprecated tools. Search volume was clearly high; existing answers were clearly weak.
Prompt 2 — The angle
“Cinematic AI requires zero cameras AND zero studio.”
- Contrarian element: most tutorials assume you film something first
- Not for: people who want fast/easy stock-footage AI workflows
- Format: case study + tutorial hybrid
Prompt 3 — The hook
“I made 1.4 million views with zero cameras.”
— The hook that beat 10 candidatesSingle specific claim with a specific number. No questions, no “in this video” filler.
Prompt 4 — The visual
Empty cinematic film set at dawn, no equipment visible, single director's chair in center frame, soft fog, golden hour through windows, shot wide angle.
That visual became the hero frame at second 0.
Step 2 — Writing the spoken script (20 minutes)
The constraints I always apply to a 60-second reel script:
- Max 90 spoken seconds (gives 50-60 sec after edit cuts)
- First 12 words = the hook
- Last 8 seconds = a takeaway statement
- No “in this video,” no “let me show you”
- One single idea
Here’s the script I shipped (90 seconds spoken):
I made 1.4 million views with zero cameras. [pause 0.5s] Three years ago I uploaded my first AI video. Got 47 views. Today my channel sits at 1.4 million on a single reel. The gap between video #43 and #178 wasn't talent. Wasn't gear. Wasn't some prompt nobody else knows. It was one decision I kept making wrong for 8 months. I was treating the AI as the product. [pause 0.5s] The AI is the brush. The product — the thing that makes a stranger watch 47 seconds of your video at 2 AM — is a perspective. For 8 months I asked: "what looks cool?" None of those videos found their audience. Then I started asking: "what question is my audience already typing into TikTok at 1 AM?" That single shift unlocked it. The Blueprint walks through every step. Link is in my bio.
That’s 168 words. At 175 WPM (slight slow-down for emphasis pauses) = ~58 seconds spoken. After cuts: 47 seconds final.
Step 3 — Voice via ElevenLabs (15 minutes)
I used my custom-cloned voice on ElevenLabs (a one-time 10-minute training I did months earlier, reusable forever).
| ElevenLabs setting | My value (viral short-form) |
|---|---|
| Model | V3 |
| Stability | 35% |
| Similarity Boost | 78% |
| Style | 20% |
| Speaker Boost | ON |
| Speed | 0.96 |
Why these settings work: defaults are tuned for enterprise narration, NOT viral short-form. These tuned-down values introduce natural variation.
Total credit cost: ~870 characters of generation = ~$0.40 on Creator plan amortized.
Step 4 — Visual generation (45 minutes)
Six visual beats matched to script. Each one selected for what the tool does best.
Beat 1 (0:00-0:08) — The empty film set
Tool: MidJourney V7
Cinematic film set at dawn, completely empty, no cameras or equipment visible, single director's chair in center frame, soft golden fog, golden hour light streaming through tall windows, shot wide angle on 35mm, atmospheric, magazine cover quality --ar 9:16 --v 7
Generated 4 variations. Picked the one with the strongest negative space.
Beat 2 (0:08-0:14) — The 47-view era
Tool: Higgsfield Soul — Me (Soul-trained avatar) sitting at a desk looking at a laptop screen showing “47 views.” Frustrated expression. I used Soul because I needed character consistency across multiple shots and Soul is best in class for that.
Beat 3 (0:14-0:22) — The “1.4M” reveal
Tool: MidJourney
Glowing neon number "1.4M" floating in dark space, cyberpunk aesthetic, depth and atmosphere, sharp focus on numbers, soft glow, shot on medium format, cinematic 9:16 vertical --ar 9:16 --v 7
This was the most-screenshotted frame from the reel. Singular focus on the number.
Beat 4 (0:22-0:30) — Brush metaphor
Tool: Kling 3.0 — close-up of a paintbrush dragging across a digital canvas, paint becoming a video frame. Kling won here because of motion physics — the paint had to flow right. Veo would’ve been slower; MidJourney can’t do motion at all.
Beat 5 (0:30-0:42) — Question being typed
Tool: MidJourney + simple text overlay — a phone screen at 1 AM showing a TikTok search bar with someone typing a question. Atmospheric blue glow.
Beat 6 (0:42-0:47) — Blueprint reveal
Tool: MidJourney — The Blueprint PDF cover shot as if held in hand. Tight crop. Final cut-to-bio CTA.
Step 5 — Editing in CapCut (30 minutes)
Rough cut (~15 min)
Dropped all 6 visual beats into CapCut in script order. Synced to voice tracks.
First 3 seconds polish (~10 min)
This is where I spend disproportionate time. The opening shot needed:
- A subtle camera move (zoom-in 102%)
- A music sting at 0:00 to match the hook word “1.4 million”
- A flash transition at second 0:01
This is the part most creators rush. Don’t.
— First 3 seconds = the algorithm testCaptions + music + end frame (~5 min)
- Auto-captions in CapCut, customized to white bold with subtle drop shadow
- Music bed: 70% voice / 30% ambient cinematic track at -18dB
- End frame: Blueprint cover + bio link CTA, 4 seconds hold
Step 6 — Distribution (10 minutes)
Exported 1080×1920 (9:16 vertical) at H.264. Uploaded to:
- YouTube Shorts — title: “I Made 1.4M Views With Zero Cameras”
- Instagram Reels — same caption, same hashtags
- TikTok — slightly different opening text overlay
Same export, three platforms. The “edit once, post three times” rule.
Why this reel hit 1.4M
Post-mortem on what actually worked:
- The question was already validated. I didn’t guess. I FOUND that “how to make cinematic AI without uncanny” was asked weekly with bad answers.
- The first 3 seconds did their job. Empty film set + bold “1.4 million” claim = pattern interrupt.
- The visual concept contradicted expectations. Cinema = cameras. Mine = no cameras visible. Contradiction earned the first second of attention.
- The script was tight. 168 words for 47 spoken seconds. No filler.
- The CTA was specific. “Link in bio” is generic. Mine: “The Blueprint walks through every step. Link is in my bio.”
What I’d do differently next time
- More retention re-hooks. Around seconds 22-25 the retention curve has a slight dip. Adding a verbal “here’s the part most creators miss” would tighten retention.
- Use Veo 3.1 for Beat 5. Kling did fine but Veo handles “phone screen close-up” with native audio better.
- A/B test thumbnail crops. I shipped one frame. Could’ve tested the empty film set frame vs the 1.4M neon frame.
FAQ
3 hours from “I have an idea” to “uploaded to 3 platforms.” 30 of those minutes were prompt-framework. The rest was execution.
Roughly $3.20 in credits + amortized subscription. ElevenLabs ~$0.40 voice. MidJourney ~$0.80 in fast hours. Higgsfield ~$1.50 for Soul + Kling beat 4. CapCut Pro amortized $0.50.
The 30-minute framework runs the same for beginners. Beat 4 (Kling motion) would be harder without practice. Beat 2 (Soul avatar) requires training your own avatar first. Realistic beginner version: same script, MidJourney still frames only, free ElevenLabs tier. Total time: 4-5 hours. Total cost: $0-10.
MidJourney is best for static frames. Higgsfield’s Soul model is best for character consistency. Different jobs, different tools.
I do — that’s my current cadence. 3 hours × 12 = 36 hours/month of focused production work. Roughly half a part-time job’s hours.
🎁 Show Your Work — Steal My Prompt
Want my 1.4M-view AI thumbnail with YOUR face on it? Here's the exact prompt + 4-step recipe to recreate it in 2 minutes using GPT Image 2 medium 2k via Higgsfield.
EXTREME LOW ANGLE HERO SHOT, 16:9 cinematic thumbnail. SUBJECT: same face from reference photos (face fidelity critical). OUTFIT: oversized Moncler Maya black down jacket fur-lined hood, Dior CD diamond crewneck, Gucci GG monogram silk scarf. JEWELRY: stacked Cuban link gold chains + Versace Medusa medallion swinging at camera, Audemars Piguet iced-out diamond chain, Richard Mille RM 11-03 skeleton dial watch on wrist. POSE: confident hard, hand on chain pulling toward camera, looking DOWN at viewer, slight smirk. SETTING: dimly-lit video editing studio, monitors mounted high glowing with MidJourney + ElevenLabs + Kling UIs. LIGHTING: dramatic underlighting neon purple #6E2BFF, teal #00E5C7 rim, golden warm spot on face, cinematic film noir. TEXT BAKED: top-left massive 'I MADE 1.4M VIEWS' bold black sans-serif with gold drop shadow, below white 'AI REEL SYSTEM', bottom-right small white 'MICHYDEV.COM'. STYLE: hyperrealistic GQ x Hypebeast hip-hop editorial, sharp focus, shallow DOF, cinematic grain.
- Pick your character reference. Use 3 well-lit frontal photos of yourself. Upload them to Higgsfield as media references — face fidelity needs 3 angles.
- Open Higgsfield → New Image → GPT Image 2. Set resolution = 2k, quality = medium, aspect ratio = 16:9. Paste the prompt above.
- Attach your 3 reference photos. This is the difference between “AI random face” and “actually YOU”. Skip this and the model invents a stranger.
- Generate, iterate angle. ~30-50s and ~10 credits per shot. Always vary the camera angle (Low Angle, Over-Shoulder, Dutch Tilt, Top-Down) — never repeat the same shot.
Want video like this for your business?
Tell me about your project. I usually reply in under 4 hours.
Get in touch