Faceless Short Video Pipeline — Chat, Stills & TTS
Build faceless YouTube Shorts, TikTok, and Reels with ForgeEcho—AI Chat scripts, 9:16 stills, ElevenLabs voiceover, CapCut assembly. No native video model required.
What “faceless” means here
A faceless short is a 15–45 second vertical video that sells or teaches without showing your face: product stills or B-roll frames + captions + TTS narration. ForgeEcho does not generate native video clips (no Sora/Veo inside the product). The production stack is:
AI Chat (script + cover prompts)
→ AI Image (9:16 key frames)
→ AI Voice (spoken track)
→ CapCut / Premiere / Descript (assemble + captions)That stack is enough for product explainers, UGC-style ads, and listicle Shorts—and it stays inside one credit system until the editor step.
Deliverable brief (fill before you open Chat)
| Field | Example |
|---|---|
| Channel | TikTok / Reels / YouTube Shorts |
| Length | 18–22 seconds |
| Offer | 30ml serum, “calmer skin in 7 days” claim (keep claims compliant) |
| Visual style | UGC handheld OR clean studio—pick one |
| Voice persona | Skeptical friend / calm expert / busy parent |
| CTA | “Link in bio” / “Shop the serum” |
One persona + one visual style per clip. Mixing “luxury studio + meme UGC voice” in the same 20 seconds usually underperforms.
End-to-end workflow (one clip, ~45–60 min)
1. Script in AI Chat (~8 min, ~1–1.5 credits)
Paste:
Write a spoken 18-second faceless ad script (~35–45 English words).
Structure: Hook (first 8–12 words) → One benefit → Soft CTA.
Persona: [busy parent / skeptic / enthusiast].
No parentheticals like (laughs). Short sentences. Brand name once.
Give 2 variants.Pick the variant you can read aloud without stumbling. Ask Chat to cut any line longer than ~14 words.
2. Cover + cutaway prompts in AI Chat (~5 min, ~1 credit)
Ask for two image prompts:
- Hook cover — product hero, 9:16, top third low-texture for captions
- Cutaway — ingredient/material close-up OR lifestyle context, same palette
Keep “no text in image, no watermark, no face” unless you intentionally want a face (then this is no longer a faceless workflow).
3. AI Image: screen then finalize (~20 min)
| Pass | Model | Size | Count |
|---|---|---|---|
| Screen | nano-banana-fast | 1K | 3–4 per prompt |
| Final | nano-banana-2 | 2K | 1–2 winners |
Upload a product reference when pack shape/label must stay accurate. Score at phone width: is the product readable in the first 3 seconds?
Credits: roughly 3×4 + 4×2 ≈ 20 for two prompt families (adjust if you screen less).
4. AI Voice: sample then full (~10 min)
- Paste the winning script
- Generate two voices at 10–15 seconds (or the full short script if already ≤45 words)
- Check: hook emphasis, brand name pronunciation, speed 0.95–1.05
Billing: 1 credit per 500 characters (short scripts usually 1 credit each). Details: AI Voice.
5. Assemble outside ForgeEcho (~15–20 min)
Recommended editor settings:
| Layer | Tip |
|---|---|
| Timeline | Cover 0–3s → cutaway → back to product on CTA |
| Captions | Burn in large type; keep within center-safe zone |
| BGM | −18 to −24 dB under voice |
| Export | 1080×1920, 30fps, H.264 |
Optional: 3–5s motion on stills via your editor’s Ken Burns / external image-to-video—after the still and VO are locked.
Three proven clip formats
Format A — Problem → Product → CTA
| Beat | Visual | VO |
|---|---|---|
| 0–3s | Frustrated lifestyle still (no face) or product in messy context | Hook question |
| 3–12s | Clean product hero | Benefit |
| 12–18s | Product + empty caption band | CTA |
Format B — Myth → Proof texture → CTA
| Beat | Visual | VO |
|---|---|---|
| 0–4s | Bold product close-up | “Stop doing X…” |
| 4–14s | Ingredient / material macro | One proof point |
| 14–20s | Pack shot | CTA |
Format C — Listicle (3 tips)
Generate three cutaways with the same palette; VO lists tip 1/2/3 in under 30s. Cap Chat scripts at ~55 words.
Weekly batch (9 clips) without melting credits
| Day | Action |
|---|---|
| Mon | Chat: 3 personas × 3 scripts; lock VO personas |
| Tue | Image: all covers at 1K; promote 9 finals at 2K Wed |
| Wed | Voice samples → finals |
| Thu | Edit + captions |
| Fri | Publish 3; hold 6 for fatigue swaps |
Budget band inside ForgeEcho: roughly 9 × 45 ≈ 400 credits if you are wasteful; with strict 1K→2K funnels closer to 280–320. Use Credit Budget & Model Selection.
Quality checklist before publish
- Hook readable without sound (caption + cover)
- Product recognizable in first frame
- VO starts within 0.3s of picture
- No garbled “AI text” on pack
- Claims match your market’s advertising rules
- Winning script + prompts saved to Prompt Library
Common mistakes
- Writing scripts for reading, not speaking
- Generating 4K covers for a 1080×1920 export
- Three visual styles in one 15-second clip
- Expecting ForgeEcho to output a finished MP4 (it won’t—by design)
FAQ
Can I use one voice across a whole channel?
Yes—lock 1–2 voices per brand for recognition. Swap scripts, not voices, during fatigue tests.
Do I need lifestyle footage?
No. Strong product stills + captions + clear VO outperform random stock B-roll for many DTC offers.
Relation to batch social guide?
Social Media Batch Creative scales templates across a calendar; this guide is the single-clip assembly SOP.