Create a story video from text

6 min readeditorstorytext-to-videottsvoiceoverb-rollgrowth

Type a topic (or paste a full script) and kuhukoo writes, voices, and assembles a narrated story video for you. Under the hood we chain Gemini for the script, Pixabay for b-roll, and a self-hosted Piper voice for the narration — all inside your account, no third-party service holds your text.

When you see the button

Composer with NO media attached → 📖✨ Create story video.

You don't attach anything first — this workflow CREATES the media. Once attached-media appears in the composer the button hides (use ✨ Auto-create video if you have your own assets, or ✂️✨ Auto-clip into shorts if you have a long video).

Tier requirement

Growth+ feature (same tier as Auto-clip and vision AI). Free and Starter tenants see the button but any click surfaces a 402 tier_gated upgrade prompt. Upgrade in Settings → Billing.

Step 1 · Settings

  • Topic — one line describing the video. Gemini writes the script + b-roll queries + optional on-screen highlights sized for your target duration at ~150 wpm.
  • Paste script (verbatim) — toggle if you already wrote your narration. We split it into scenes but keep every word exactly as you wrote it (a fidelity guard rejects paraphrased Gemini responses and falls back to a local sentence splitter).
  • Duration — 30–120 s.
  • Tone — informative / casual / energetic / cinematic / witty. Shapes the script only; TTS voice is picked separately.
  • Aspect — 9:16 (Reels/Shorts default), 1:1, 4:5, 16:9.
  • Voice — three commercially-licensed voices:
  • Amy (US · warm female, MIT-licensed LJSpeech corpus)
  • Hannah (US · newscast female, BSD-2 HiFi Captain corpus)
  • Henry (US · steady male, BSD-2 HiFi Captain corpus)

Hit 🔊 Sample to hear a phrase before choosing.

Step 2 · Review scenes

Every scene has:

  • Narration — edit inline. In topic mode you can rephrase freely; in paste-script mode you're editing your own original words.
  • b-roll query — 2–4 keywords the Pixabay stock library searches. If the auto-pick isn't right, change the query and the composer re-fetches when you hit Assemble.
  • On-screen text — optional short overlay for that scene.
  • Reorder (↑ / ↓) and Delete per scene.

The title + total-duration estimate at the top updates as you edit.

Step 3 · Assemble

Progress phases:

  1. Writing script — one Gemini call, cached-JSON parsed, fidelity-checked (paste-script mode).
  2. Fetching visuals — per scene: Pixabay search → attach the top hit to your tenant. Empty results fall back to a solid-color still card with your on-screen text.
  3. Recording voiceover — per scene: one Piper TTS call, with a single retry on transient errors (429/503/504). A scene that still fails becomes a silent slideshow entry with captions — the story still ships and you can hand-fix in the editor.
  4. Assembling — all pieces compose into a VideoEditSession the editor opens.

What ships in the assembled session

  • 9:16 (or your chosen aspect), one segment per scene, Ken Burns motion on stills.
  • Voiceover lane (E17b) — one track per narrated scene, positioned at that scene's start time.
  • Music bed — optional; when attached, it auto-ducks under voiceover via the E15 sidechain follower.
  • Captions — one text block per sentence, distributed across the scene span proportionally to word count. Honest limit: these are approximate word timings, not per-word from an ASR pass — you can hand-tune every block in the Captions tool.
  • On-screen text — top-third overlays with visibility windows matching the scene.
  • audioMix.duckOriginalUnderVoiceover on so any b-roll audio dips under the narration.

The composer opens the editor with the assembled session. Save it → the render pipeline (ffmpeg.wasm) exports a single mp4 → draft Post created with the story title as the caption seed.

Guards

  • Max 12 scenes — long stories are capped both at parse time and in the assembler.
  • Empty b-roll — solid-color still card fallback.
  • TTS failure — one retry, then silent scene (you can add voiceover manually in the editor).
  • Low-power devices — MediaPipe smart-crop from E18 also runs post-assemble if you're on 9:16 aspect and the session ever needs re-cropping; the same low-power auto-default applies.

Voice licences (audit-ready)

Voice Model licence Corpus source
Amy MIT LJSpeech (public-domain-like)
Hannah BSD-2-Clause HiFi Captain
Henry BSD-2-Clause HiFi Captain

All commercially usable. Lessac was considered and rejected — its Blizzard 2013 dataset is research-only.

Feedback

Multi-language voices (Hindi / Punjabi via Sarvam) are on the roadmap. Descript-style edit-by-transcript is next up. Report the story shapes you WISH existed so we can pick the next TTS voice + prompt template: /help/report-bug.

Was this helpful?

Still stuck? Contact support — we reply within one business day.