ホームに戻る

Faceless Short Video Pipeline — Chat, Stills & TTS

Build faceless YouTube Shorts, TikTok, and Reels with ForgeEcho—AI Chat scripts, 9:16 stills, ElevenLabs voiceover, CapCut assembly. No native video model required.

What “faceless” means here

A faceless short is a 15–45 second vertical video that sells or teaches without showing your face: product stills or B-roll frames + captions + TTS narration. ForgeEcho does not generate native video clips (no Sora/Veo inside the product). The production stack is:

AI Chat (script + cover prompts)
  → AI Image (9:16 key frames)
  → AI Voice (spoken track)
  → CapCut / Premiere / Descript (assemble + captions)

That stack is enough for product explainers, UGC-style ads, and listicle Shorts—and it stays inside one credit system until the editor step.

Deliverable brief (fill before you open Chat)

FieldExample
ChannelTikTok / Reels / YouTube Shorts
Length18–22 seconds
Offer30ml serum, “calmer skin in 7 days” claim (keep claims compliant)
Visual styleUGC handheld OR clean studio—pick one
Voice personaSkeptical friend / calm expert / busy parent
CTA“Link in bio” / “Shop the serum”

One persona + one visual style per clip. Mixing “luxury studio + meme UGC voice” in the same 20 seconds usually underperforms.

End-to-end workflow (one clip, ~45–60 min)

1. Script in AI Chat (~8 min, ~1–1.5 credits)

Paste:

Write a spoken 18-second faceless ad script (~35–45 English words).
Structure: Hook (first 8–12 words) → One benefit → Soft CTA.
Persona: [busy parent / skeptic / enthusiast].
No parentheticals like (laughs). Short sentences. Brand name once.
Give 2 variants.

Pick the variant you can read aloud without stumbling. Ask Chat to cut any line longer than ~14 words.

2. Cover + cutaway prompts in AI Chat (~5 min, ~1 credit)

Ask for two image prompts:

  1. Hook cover — product hero, 9:16, top third low-texture for captions
  2. Cutaway — ingredient/material close-up OR lifestyle context, same palette

Keep “no text in image, no watermark, no face” unless you intentionally want a face (then this is no longer a faceless workflow).

3. AI Image: screen then finalize (~20 min)

PassModelSizeCount
Screennano-banana-fast1K3–4 per prompt
Finalnano-banana-22K1–2 winners

Upload a product reference when pack shape/label must stay accurate. Score at phone width: is the product readable in the first 3 seconds?

Credits: roughly 3×4 + 4×2 ≈ 20 for two prompt families (adjust if you screen less).

4. AI Voice: sample then full (~10 min)

  1. Paste the winning script
  2. Generate two voices at 10–15 seconds (or the full short script if already ≤45 words)
  3. Check: hook emphasis, brand name pronunciation, speed 0.95–1.05

Billing: 1 credit per 500 characters (short scripts usually 1 credit each). Details: AI Voice.

5. Assemble outside ForgeEcho (~15–20 min)

Recommended editor settings:

LayerTip
TimelineCover 0–3s → cutaway → back to product on CTA
CaptionsBurn in large type; keep within center-safe zone
BGM−18 to −24 dB under voice
Export1080×1920, 30fps, H.264

Optional: 3–5s motion on stills via your editor’s Ken Burns / external image-to-video—after the still and VO are locked.

Three proven clip formats

Format A — Problem → Product → CTA

BeatVisualVO
0–3sFrustrated lifestyle still (no face) or product in messy contextHook question
3–12sClean product heroBenefit
12–18sProduct + empty caption bandCTA

Format B — Myth → Proof texture → CTA

BeatVisualVO
0–4sBold product close-up“Stop doing X…”
4–14sIngredient / material macroOne proof point
14–20sPack shotCTA

Format C — Listicle (3 tips)

Generate three cutaways with the same palette; VO lists tip 1/2/3 in under 30s. Cap Chat scripts at ~55 words.

Weekly batch (9 clips) without melting credits

DayAction
MonChat: 3 personas × 3 scripts; lock VO personas
TueImage: all covers at 1K; promote 9 finals at 2K Wed
WedVoice samples → finals
ThuEdit + captions
FriPublish 3; hold 6 for fatigue swaps

Budget band inside ForgeEcho: roughly 9 × 45 ≈ 400 credits if you are wasteful; with strict 1K→2K funnels closer to 280–320. Use Credit Budget & Model Selection.

Quality checklist before publish

  • Hook readable without sound (caption + cover)
  • Product recognizable in first frame
  • VO starts within 0.3s of picture
  • No garbled “AI text” on pack
  • Claims match your market’s advertising rules
  • Winning script + prompts saved to Prompt Library

Common mistakes

  • Writing scripts for reading, not speaking
  • Generating 4K covers for a 1080×1920 export
  • Three visual styles in one 15-second clip
  • Expecting ForgeEcho to output a finished MP4 (it won’t—by design)

FAQ

Can I use one voice across a whole channel?
Yes—lock 1–2 voices per brand for recognition. Swap scripts, not voices, during fatigue tests.

Do I need lifestyle footage?
No. Strong product stills + captions + clear VO outperform random stock B-roll for many DTC offers.

Relation to batch social guide?
Social Media Batch Creative scales templates across a calendar; this guide is the single-clip assembly SOP.

Related Guides

  • AI Chat
  • AI Voice
  • Social Media Batch Creative
  • Credit Budget & Model Selection
What “faceless” means hereDeliverable brief (fill before you open Chat)End-to-end workflow (one clip, ~45–60 min)1. Script in AI Chat (~8 min, ~1–1.5 credits)2. Cover + cutaway prompts in AI Chat (~5 min, ~1 credit)3. AI Image: screen then finalize (~20 min)4. AI Voice: sample then full (~10 min)5. Assemble outside ForgeEcho (~15–20 min)Three proven clip formatsFormat A — Problem → Product → CTAFormat B — Myth → Proof texture → CTAFormat C — Listicle (3 tips)Weekly batch (9 clips) without melting creditsQuality checklist before publishCommon mistakesFAQRelated Guides