Skip to main content
1

Upload your audio

Drop a voiceover; we transcribe it into timestamped lines.

StoryAnimator.app · illustrated story video studio Pay as you go

Turn any voiceover into an illustrated story video

Make faceless YouTube videos from a voiceover — no camera, no face needed. Drop in an audio narration or record your own. StoryAnimator transcribes it into timestamped sentences, generates one consistent image per line, and lays it onto a real editing timeline synced to the waveform — ready to export as MP4.

A voice waveform turning into an animated story video
8
Timestamped lines
0:58
Narration length
6
Image models
16:9
Export ratio
1

Upload your audio

Drop a voiceover or narration file. We run speech-to-text and split it into sentences with start and end times.

Your narration

Drop audio or video here, or click to browse

Audio (WAV · MP3 · M4A · AAC · FLAC) or video (MP4 · MOV · WEBM) — up to ~35 min. A video's soundtrack is used automatically.

Auto-detect language · timestamped sentences
source.wav
No file yet
Waiting for audio…
Waiting
Upload an audio file to begin transcription.
2

Review the transcript

Every sentence becomes a shot. Hover or click a row to highlight it — these timings drive image durations on the timeline.

Transcript & shots
transcript.txtread-only
shots.list0 shots

How many images?

drag or step — each shot becomes one generated image

0
fewer more
Story style

Pick one art style and set three guardrails — they apply to every image so the whole video feels like one film.

Art style
Pick a preset or describe your own — this single choice drives the whole look.
Applies to the cast and every generated frame.
Visual styleMood, medium, palette, lighting
Character bibleFixed designs reused every shot
Negative promptWhat to keep out of frame
Environment · optionalLeave empty — each shot follows the story. Fill only to pin one fixed setting everywhere.
Generation setup
Image modelWhich AI draws each frame
Video formatSets the shape of every image
Reference imageoptional
Upload character ref
Upload voice first to unlock this step.

Character cast

Generate a portrait of each character, refresh any you don't like, or edit a prompt with the pencil — the approved cast locks character consistency.

5

Results gallery

• Each generated 16:9 frame shows its index, filename, and the prompt that produced it.

• Export the whole job as a ZIP with manifest.

• Animating shots is optional — skip it and click Continue to export your video from the still images.

storyboardno job yet
Edit any shot's image prompt, then hit Generate. Re-roll individual frames anytime after.
Tip: click a frame to select it · shift-click for a range · ⌘/Ctrl-click to add · right-click for Delete / Cut / Copy / Paste (paste asks before or after).
estimate
Video modelHow each still is animated into motion
How to animateDefault motion & feel for every shot — per-shot “🎬 Motion” overrides it
Image quality
Nano Banana keeps characters consistent, accepts a reference image, and follows the per-sentence prompt.
$0.0000
0 frames @ $0.0000 · estimated total
Transcribe audio before generating.
6

Video preview

Every frame laid end to end, each clip stretched to its sentence duration over the narration waveform. Scrub, preview, and export to MP4.

sequence.editno sequence
SHOT 01
PREVIEW
0:00.0 / 0:00.0
Shot timeline
0:000:00
Narration volume
Background musicNo music selected
Upload music you own, or licensed / Creative Commons music with attribution.
Subtitles
Subtitle styleKaraoke styles sync word-by-word to the voice (preview & full export)
Captions come from your timestamped transcript and stay in sync with each shot — no manual timing needed.
Import
Rebuild a video from existing assets — no regeneration. Three ways in:
① Video — split into frames + audiomp4, mov, webm
Splits the video into shot frames and lays its audio underneath. Leave Frames blank to auto-pick (~1 every 3s).
② Frames onlytimestamp filenames (e.g. 049.png)
Each frame is timed from its filename (the timestamp). Pair with a soundtrack below.
③ Soundtrack / audiomp3, wav, m4a…
Sits in the audio track underneath the frames. No re-transcription.
Reconstruct the project (reverse pipeline)subtitles + name + style
Transcribes the imported soundtrack, writes a caption onto each frame by its timing, and drafts the project name (if none), visual style, character bible, intro/outro & thumbnail — the reverse of the normal pipeline.
Export qualityformat set in Look
~ size shown here
Render
Ready · records in real time with audio

YouTube details

Title, description, and tags — auto-written from your narration when the video is ready. Edit anything before you post.

▣ PUBLISH TO YOUTUBE
publish.meta
Title 0 / 100
Tags comma sep.
Description
These fill in automatically once your video finishes exporting — or click ✨ Generate with AI to write now.
Open YouTube upload ↗ Export your MP4 in the Editor, click Open YouTube upload, drop in the file, then paste your copied title / description / tags.
thumbnail.genfirst free · regen 1 credit
Click Generate to preview
Overlay text — ≤4 bold words
Image prompt — no text in image
1280 × 720 · ready to upload to YouTube
$

Credits and usage

Transcript and ZIP exports are free. Image generation and video export use credits.

total.bill
Transcription$0.0000
Character portraits
Image generation$0.0000
Video render$0.0000
Total spent this session
$0.0000
Regular users see credits only. Admins see internal USD cost for margin tracking.

History

Your saved projects. Re-open one to view its frames; re-rolling a shot replaces the stored image with the newest version.

history.log
Projects are saved automatically when you generate images — and updated whenever you re-roll a still or re-animate a shot, so this always shows your latest version.
Done