Audio & Voice Pipeline

How dialog lines become audio -- TTS, voice changer, recording, and merged character tracks.

Ordinary Animator has a full dialog audio pipeline: from dialog lines written in the screenplay, through voice synthesis or recording, to per-character merged audio tracks that play back on the timeline. The episode's whole spoken track lives on the Dialogue Audio view, where you can generate it in one press and work through it line by line. This page covers the one-press surface; for the small operations underneath it -- and the recipes the platform builds from them -- see Building Dialogue from Operations.

Dialog lines

Dialog is written in the screenplay (Fountain format) and parsed into the scene's dialog structure. Each line is associated with a character and a shot. The Dialog tab in a scene or shot shows all lines in order.

Every line can independently have:

  • Generated TTS audio
  • A recorded audio file (uploaded or captured)
  • A voice-changer version -- a recording re-voiced as the character

Voice configuration

Each character has a Voice configuration. You select which TTS engine to use and which voice within that engine:

Engine Notes
Google Fast, consistent, and cheap -- the recommended default for a first pass
ElevenLabs High-quality synthesis with many voice options; requires API key
Typecast Alternative TTS engine
ComfyUI Run text-to-speech or a voice changer through your own ComfyUI workflows

A character can have more than one voice set up at once (for example Google for a quick first cut, plus a voice changer for hero lines), and you star one as the draft voice. Set this up on the character's Voice page. The one-press generation uses each character's draft voice, so there's nothing to pick at run time.

Generating dialogue audio

You rarely generate audio one line at a time. The fastest path is one press for everything:

  • On the Dialogue Audio step of the episode plan, Generate all dialogue audio works through every shot in the episode: it creates the audio for each line and lays each character's lines out with the right timing for the shot, ready for animation. Each scene also has its own Generate this scene button, and the same button appears on a scene's page. (Dialogue you haven't placed in a shot yet is left until you add it to one -- you'll be told if there's any.)
  • The button just runs -- it uses each character's starred draft voice, so there's no per-run prompt. Characters whose draft voice is a voice changer, or who have no voice set up, are listed for you rather than synthesised.
  • It only fills in text-to-speech. Lines you intend to record, run through a voice changer, or let a video model voice are left untouched and listed afterwards so you know what still needs a human step -- they are never treated as errors.
  • Re-running is safe: lines that are already up to date are left alone; only new or changed lines are regenerated.

This is built for an audio-first preview: generate the whole episode's voice track in one click, listen end to end, then upgrade individual lines (record a performance, clone a voice) where it matters.

For per-line control -- alternates, recording, voice changing, picking the preferred take -- open a shot's media gallery and use the Takes panel. You can preview each clip inline, regenerate, and choose between takes.

Reviewing and choosing takes

A line can end up with several takes -- a quick first pass, a higher-quality alternate, a recorded performance, a voice-changer version. You never have to choose: generating a line automatically marks the newest take as the one that's used, so there's always a usable voice. Choosing is optional polish.

Picking is done right where the lines are: on the Dialogue rendered / recorded step, each line shows how many takes it has. Expand a line to play each take and choose the one you want -- or generate, record, or clone another take on the spot. Picking a take is the only choice you make per line.

Because a shot's combined per-character audio is built from the chosen takes, changing a choice puts the affected shots' combined audio out of date. The same step shows a single Re-merge updated shots button whenever any shot is out of date, so rebuilding the timing track to match your picks is one click.

Previewing the episode

The Dialogue Audio step ends with Preview the episode -- the quickest way to hear how everything fits together before you invest in final renders.

  • Prepare preview audio generates any missing dialogue and merges each shot's audio in one press (the same cache-aware run as Generate all dialogue audio).
  • A per-shot checklist shows which shots are ready and which still need audio.
  • Preview episode on the timeline opens the timeline, where pressing Play lays the dialogue over whatever media each shot has -- a storyboard still or a video -- with any music cues over the top. You don't need a finished video to hear the episode: a storyboard with voices is enough to judge pacing and performance.

If a shot's video already includes its dialogue (a finished lip-synced render), the preview uses the video's own sound instead of laying the dialogue over it, so you never hear the voices twice. A video without baked-in dialogue -- for example an early proof-of-concept clip -- simply gets the dialogue played over it.

Shots you haven't given artwork yet don't show a black screen: the preview stands in a starred image -- the speaking character's, or failing that the scene's location -- so you can still follow the episode end to end. A shot also holds on screen at least as long as its dialogue, so a line is never cut off.

Recording and uploads

If you prefer a human voice or want to capture reference audio:

  • Record -- capture audio directly in the browser
  • Upload -- attach an existing audio file

Uploaded and recorded files go through the same curating flow as TTS: you preview, approve, or replace.

Voice changer

A voice changer turns a recorded performance into a character's voice, keeping the original timing and delivery. Set one up on the character's Voice page (ElevenLabs speech-to-speech or a ComfyUI workflow) and attach reference voices as the target timbre. This lets you perform a line yourself -- or in any voice -- and have it come out sounding like the character. It's the path to use when TTS isn't enough and you want a real performance in the character's voice.

Merged audio tracks

When all dialog lines for a scene are generated or recorded, you can Merge audio per character. The merge produces a single audio track per character, with silence between spoken lines, timed to align with shot durations.

The merged track is what the timeline renderer uses when assembling the final video -- it's mixed across all characters in the scene.

The render pipeline

See Shots for how rendered video clips and audio tracks are combined in the shot and scene timeline render.