From Storyboard to Video
How a planned shot becomes a finished clip -- storyboards, the media that feeds the models, and how recipes and the AI carry the variable work.
There is a gap between a shot on paper -- "Mira turns and sees the door" -- and a finished video clip. Closing that gap is the heart of production, and there is never just one way across. One shot might start from a hand-drawn frame; another from a single line of text; another from a reference video you crop to length. Ordinary Animator does not try to flatten that variety into one rigid procedure. It gives you two clear goals, lets recipes and the AI carry the variable work in between, and stays out of your way if you'd rather do it by hand.
Two goals, not many phases
Shot work has two deliverables, and each phase is named for the media you end up with:
- The Boards -- a quick storyboard image for every shot.
- The Shoot -- a finished video clip for every shot.
That's deliberately small. Everything between a storyboard and a clip is too variable to pin into fixed steps, so we don't -- recipes and the AI handle it. The pipeline tells you what you're trying to have; it doesn't dictate how you get there.
The Boards: the storyboard
A storyboard is a quick image -- a preview. It can be a hand sketch, a fast AI pass, or one frame pulled from a multi-angle reference. Fidelity is not the point; blocking is.
Blocking is working out the staging of the shot: who is in frame and where they stand, how they're posed and what they do, where the camera sits and whether it moves, and what the audience's eye is drawn to. Borrowed from theatre and animation, it's the cheap place to get staging right -- redraw a rough frame in seconds instead of paying to regenerate a finished clip.
As you board a shot you also capture details about it -- shot type, who's in it, camera angle, motion, rough duration. This isn't busywork or a form to grind through: those details are what the generation models and the AI read later to produce the right image and video. The AI helps fill them in and flags what's missing, so you're making creative decisions, not data entry.
Media, tagged by purpose
The concrete thing that flows through production is media, tagged by its purpose. The media gallery shows every image, clip, and audio file; each one carries a tag for what it's for:
- storyboard -- a quick preview frame
- start frame -- a high-quality still used to seed image-to-video
- plate -- a background image to composite onto
- reference -- a character, style, or location image the model should match
- driving clip -- a video whose motion or expression is copied
- audio -- the dialogue and music you've already produced
Recipes read tagged media and write new tagged media. The same purpose can be reached many ways -- a start frame might come from a sketch, from a text prompt, or from a posed 3D figure -- so the tag is the stable, useful thing, and the route to it is not worth naming.
The Shoot: producing the clip
The second goal is a video clip per shot, and this is a single production step -- because the path is recipe-shaped and varies from shot to shot:
- Text-to-video -- just a prompt; nothing to prepare.
- Image-to-video -- make a high-quality start frame first, then animate it.
- Reference-to-video -- no start frame at all; a rough blocking clip and some reference images set the shot.
- Video-to-video -- crop a driving clip to length, gather reference images, generate.
- Static + effects -- a slow camera move over a still; no character motion to get wrong.
Sometimes you build a polished still before generating; sometimes you go straight to the clip. That variability is exactly why this isn't carved into fixed sub-steps -- the recipe decides what's needed, and the recipe is the concrete, repeatable thing that does the work.
The AI rides alongside. It looks at how the shot was blocked and suggests an approach in plain language -- "one good route here is a hand sketch into an image model for the start frame, then image-to-video" -- and hands you the recipe to run. Before you spend generation time, it checks you have the right inputs: "this model needs a clean, face-forward start frame, and yours is a wide shot -- consider a static shot with a push-in instead." That suitability check is cheap and happens before you pay for a render.
Reference-to-video: directing the framing without a start frame
Image-to-video inherits its plate exactly. The first frame of the clip is the start frame you fed it, so if a character sits too small or too far left in that picture, no prompt fixes it -- you have to go back and remake the still. That makes composition un-directable at the video step, which is the one place you can actually see the motion you are composing for.
Blocking video to shot is the route that has no start frame. You give it two kinds of reference and it builds the clip from them:
- A reference video -- a rough blocking pass. It supplies the camera move, the staging, how big each character reads in frame, and the timing. It does not have to look finished; grey blocks moving through a 3D scene are enough, because you are directing the staging, not the picture.
- Reference images -- character sheets, a location plate, a style key. These supply identity and look. Up to three slots.
Nothing anchors the first frame, so reframing a shot means redoing the blocking pass rather than regenerating a plate and hoping.
Name every reference in the prompt
This is the part that decides whether it works. The model addresses your references by the order the slots are filled, not by slot name, and it only uses a reference deliberately if the prompt names it:
- the video is
<Video 1> - the images are
<Picture 1>,<Picture 2>,<Picture 3>in slot order
So write the prompt around those tags and say what each one drives:
Follow the camera move and the staging of
<Video 1>exactly. The tall figure on the left is the woman from<Picture 1>; the seated figure is the man from<Picture 2>. Keep their faces, hair and clothing identical to the references. Warm late-afternoon light, handheld feel, no cuts.
Output is very sensitive to this wording. If you move a picture from slot 2 to slot 3, update the prompt to match, or the description now points at the wrong reference.
There is no negative prompt here
This model has no negative prompt -- there is nothing for one to act on, so any text you put in a negative field is discarded without a warning. Write every exclusion as prose in the positive prompt instead: "the camera holds its distance and never pushes in" rather than a negative listing "camera push-in". A shot that keeps drifting toward camera despite a negative telling it not to is this, every time.
Joining clips into one
Video models give you a few seconds at a time, so a shot that runs longer arrives as a run of short clips rather than one. Join video clips puts them back together without leaving Ordinary Animator: drop your clips into the slots -- clip 1 plays first, clip 9 plays last -- and run the step. You get one video back in the same gallery, ready to star or feed into a later step. Slots you leave empty are skipped, and one clip in is simply one clip out.
The option that matters is "Drop each clip's last frame", and it is on by default.
When you make a run by chaining -- generating each clip from the last frame of the one before, so the motion continues -- every clip ends on the still the next one starts on. Join them with every frame kept and that shared still plays twice at each seam, which reads as a small visible hitch right at each join. The option removes exactly one copy: the final frame of every clip except the last, which keeps its own ending because nothing follows it. Turn the option off when your clips are genuinely separate pieces -- alternate takes, unrelated shots -- and nothing is shared at the seams.
Clips do not have to match. They may come from different models at different frame rates, and they are re-encoded into a single video rather than glued together as-is, so mismatched sources still play properly. A clip shaped differently from the first one is letterboxed into the first clip's frame, never cropped -- black bars are visible and you can simply redo the join, whereas a silent crop quietly throws away the edges of your picture. Sound is kept where clips have it. If a clip you dropped in cannot be read, the whole join fails and names the slot rather than handing back a shorter video that looks fine and has silently lost a clip.
The AI guides; you don't grind a checklist
The goal of all this captured detail is twofold -- feed the models the specifics they need, and support good creative choices through AI review. It is not a 40,000-item checklist to tick off by hand. The AI fills in what it can, reviews each shot with awareness of your project's genre and tone, and surfaces only what needs your judgement. You stay in the creative seat; the system carries the bookkeeping.
Review at the scene level
You generate clips shot by shot, but you review them by scene -- because the qualities that matter most are properties of the scene, not the single shot:
- Does the motion hold together and feel right for the style?
- Is continuity preserved across adjacent shots -- clothing, props, lighting, screen direction?
- Is the visual style consistent, and does each character stay recognisable?
- Does the sequence tell the scene clearly?
One scene-level review, with notes against each quality, replaces a long per-shot sign-off. Watch the scene's clips together, fix what doesn't hold, approve the scene.
The model knowledge base
The AI's suggestions are only as good as what it knows about each model. That accumulated wisdom -- "image-to-video with this model shines from a clean full-face frame but falls apart on wide, multi-character shots" -- is the model knowledge base: an experienced artist's instincts, written down so the AI can apply them on every shot. A new model arrives as a new entry, not a rebuild. (This is the same knowledge described in Models, Workflows & Recipes.)
It ships with the platform and the AI reads it live. Each model carries the look it produces, what it is good for, how much graphics memory it wants, and the settings worth starting from -- so a suggestion can name a concrete model and cite its settings ("20 steps at cfg 2.5") rather than gesture at a family. The same lookup also tells the AI whether that model fits your hardware, which is why it will say a model needs the cloud instead of recommending something your card cannot load.
You don't have to follow the pipeline
The two-goal flow is guidance, not a gate. If you already know exactly how you want to make a shot, open the media gallery, run the recipes you want, and select your final clip. The pipeline is there to help when you want a path through the variety -- not to stand between you and the work.