Skip to main content

The AI Video Pipeline That Actually Works: Characters, Animation, Edit

About Us
Published by yuliya.dzemidchuk
24 July 2026

AI Video Creation Pipeline

 

Introduction

AI video production has moved from a novelty to a working tool, and that matters right now because it changes what a small team can realistically ship. Work that used to require a shoot, a crew, a location and a budget can be assembled by one person in a few days, which puts video inside the normal marketing and product cycle instead of outside it: concept films that sell an idea internally before anyone commits to a real production, campaign and social creative that can be re-cut for each channel, product explainers, pitch visuals, and quick look-and-feel tests on a new direction. The point is not to replace filming, but to make the visual part of a project cheap enough to iterate on and fast enough to keep up with the rest of the work. 

So, we decided to test that end to end on a real project.

This article is the result. Instead of a general overview, we want to share what actually worked for us: the practical lifehacks we picked up, the mistakes we ran into, and the workflow we would recommend to any colleague who wants to produce video with AI. We walk through the whole thing on a real project, so you can see not just the theory but the real generations, prompts, and dead ends behind each step.

 

Where to look for inspiration

Sometimes you have a task but no idea how to approach it.

The core principle is simple: look for inspiration everywhere. Watch as much cinema, architecture, and painting as you can, and train yourself to notice everything beautiful around you. Creativity feeds on what you consume, so the more visual references you gather, the richer your own work becomes.

On the practical side, Pinterest is our go-to. You can type in almost any query and instantly get dozens of images to build from. It is a fast way to shape the visual direction before you write a single prompt.

 

Build a moodboard

Once the idea has come and you have gathered enough references, the next step is to collect it all into a moodboard. This is a must. You need the picture in front of your eyes before you start producing.

Into the moodboard we put concrete examples of the angles we wanted. For our product, for instance, we collected shots of a bottle held in the hands, so we would know exactly which framing we were aiming for later. We gather everything in FigJam, so the whole board sits in one place in front of us. Only after that do we move on to production.

moodboard

Our moodboard for the project, built in FigJam: hiking references, the autumn palette, and the exact bottle-in-hands framing we were aiming for.

Tip: A moodboard is not a mood collage of pretty pictures. It is a planning tool. Collect the exact shots and angles you will need, and your storyboard almost writes itself.

 

Step 1. Creating the character

Why Midjourney, and personalization

For the heroine we chose Midjourney, because this model handles human faces best and is also great at unusual, atmospheric locations. It is highly artistic, but that comes with a downside: it is not very obedient. Midjourney tends to add a lot of its own interpretation, so you have to run many generations before you get what you want.

One thing that matters a lot in Midjourney is setting up personalization, your profile. You pick the variations you like from the many the model offers, and over time it learns roughly what kind of result you expect. For example, we consistently choose people with realistic skin, visible pores, and a natural look, so the model starts generating more believable content for us.

 

Building the prompt through a language model

Midjourney understands prompts best in English, so to build one we use a language model, Claude or ChatGPT. We describe the character we want in plain words, and the model turns it into a proper English prompt.

In our case we wrote that we wanted a 20 year old girl with freckles and red curly hair, and specified that it should be a studio photograph on a white background, shot on an 85mm lens, and so on. Then we ran a batch of generations and picked the heroine we liked most.

Building the prompt through a language model

The heroine we picked out of the Midjourney batch: red curls, freckles, natural skin with visible texture.

 

Dressing her, and a color fix

Next we decided to dress the heroine, because it is better to build the character sheet with her already dressed. For this we switched to Nano Banana, a far more obedient model that follows the request precisely, though with less artistic flair than Midjourney. We asked it to dress her in a trekking jacket and put a backpack on her back.

Dressing her, and a color fix

Dressing the heroine in Nano Banana: the chosen portrait attached as the identity reference, with only the clothing added.

Then we noticed the orange jacket we had specified was blending in with her red hair, so we decided to change it to khaki. We wrote a short prompt along the lines of “do not change anything in this image, only change the jacket color to khaki.”

Dressing her, and a color fix

Changing the jacket to khaki in Nano Banana: a short local prompt, only the jacket colour changed and nothing else.

Dressing her, and a color fix

The result: the same heroine with the jacket now khaki, and nothing else touched.

Tip: Two things here. For a small fix like this the prompt can stay very short, you do not need to spell out the whole scene. And this is a precise local edit, not a new generation: when a frame is almost right, fix the one detail instead of starting over.

 

The 360 degree character sheet

Once she was dressed, we built her character sheet. This is the key step for keeping a character consistent. Our sheet had two parts: a set of close up portraits (front, profile, three quarter, back) and a full body turnaround showing her from head to toe from every side, with the same jacket, jeans, boots, and backpack in each view. We used the selected portrait as the identity reference and asked the model to build a professional reference sheet of the same woman: photorealistic, live action, clean white background, technical turnaround. We ran several generations and picked the one where the face stayed most consistent across all the angles.

 The 360 degree character sheet

The portrait side of the character sheet: the same face held across front, profile, three-quarter and back.

The 360 degree character sheet

The full-body turnaround: identical jacket, jeans, boots and backpack from every side.

The 360 degree character sheet

The prompt behind the reference sheet: the portrait locked as the identity reference, two rows, every angle, the same outfit in each view.

Tip: Lock down both the face and the full figure, including clothing and backpack. If you only fix the face, the model will start inventing the outfit in new scenes.

 

Step 2. Building the shots for animation

The logic of this stage is that you build the still frames first, then animate them later. So at this point we were not animating yet. We were generating the key frames for each scene, using the character sheet as a reference so the heroine stayed consistent.

 

Generating the locations

By our script the heroine wakes in a tent, walks through an autumn forest, and climbs into the mountains, so we needed three environments. For locations we again used Midjourney, because it creates beautiful, unusual environments (Nano Banana does this more simply and with less artistry). We generated a tent shot, a forest trail, and a mountain scene, running several batches each and picking the frames that worked. Across all of them we kept a single visual key: the same autumn palette, cinematic film grain, a slow luxury editorial feel, and our saved personalization profile. That way everything would sit in one style later in the edit.

Generating the locations

Establishing shot of the tent, generated in Midjourney and kept in the autumn palette.

Generating the locations

Forest-trail location grid, same visual key (raw, stylize 150, our saved profile code).

Generating the locations

Mountain location grid for the climb and the summit.

Tip: Think about the edit while you generate. Keeping one light and palette across all frames makes the montage clean, because a single frame that falls out of the color instantly shows in a cut.

 

Two ways to build a prompt

There are two main ways we create prompts. The first is describing the scene in plain words. That is how we did the mountains: we wrote that we wanted distant peaks with a trail leading toward them, a path a person could walk along.

The second way is working from a reference. For the tent shot we took a photograph we liked, dropped it into a language model, and asked it to describe the image from a director's point of view: the light, the composition, the mood. Then, based on that description, we built the prompt. So we keep two go-to methods: describe from scratch, or feed a reference and let the model analyze it first.

 

Putting the heroine into the scene

Here is a key technique in action. We had the tent frame and our heroine's character sheet, and we needed her, our specific redhead, lying in that tent. So in Nano Banana 2 we combined two images: photo 1 gave us the pose, framing, and location, and photo 2 was our heroine. We asked the model to replace the face and identity on photo 1 with the woman from photo 2, keeping her exact features (eye shape, blue-grey eyes, freckles, red curly hair), while preserving the overhead framing and the pose, and relighting the scene for dawn inside the tent.

Putting the heroine into the scene

Putting our heroine into the tent: photo 1 gives the pose and location, photo 2 the identity, plus a relight for dawn.

For most of this work we used Nano Banana 2, at the moment the most advanced version of Nano Banana. When it could not give us what we wanted, we would switch to the Pro model as an experiment, and sometimes it managed exactly the frame that Nano Banana 2 kept stumbling on.

Tip: A nuance about prompt length. For simple edits, like changing a jacket colour, a short one-line prompt is enough. But for complex actions, like a full face-and-identity swap where many features must be preserved, spell everything out: eye shape, eye color, freckles, hair, framing, lighting. The more complex the task, the more detailed and precise the prompt.

Tip: The real power of the pipeline: you do not have to generate the perfect frame in one shot. You build the composition you want, then swap in your consistent character and adjust the light and mood, without losing the heroine's identity.

 

Alternating shot sizes, and the boots detail

By this point we already had a rough picture of the shots we needed. We knew one important editing principle: you alternate shot sizes. You never run wide, wide, wide in a row, because it is boring to watch. There has to be variation, a wide shot, then a close up, then a medium. So after the establishing shots we wanted a detail shot, the heroine tying her boots, to break up the rhythm.

Here we learned why attaching the location reference matters. We wrote a simple, short prompt for a close up of the boots at the tent entrance. On our first try only the character reference attached, and the frame came out with flat daytime light, a green forest, and a generic feel that fell completely out of our film. On the second try we attached both references, the heroine and the tent location, and the frame came out with cold dawn mist, the warm glow of the lantern, and autumn leaves.

Alternating shot sizes, and the boots detail

Boots detail, first try: only the character reference attached, so the light came out flat and generic.

Alternating shot sizes, and the boots detail

Same prompt, both references attached: cold dawn mist and lantern glow, back in the film's palette.

Tip: Attach the character and the location, and check before generating that both actually loaded. Do not rely on the model to remember the atmosphere on its own. Without the location reference the light and mood drift, and the frame falls out of the overall style.

 

Matching the light, and walking into the forest

We also ran into an issue integrating the heroine into locations. At first Nano Banana 2 pasted her in badly, so she looked cut out, with lighting that did not match the scene. It stayed ugly until we added one line to the prompt: adjust her lighting and color temperature to match the location. After that it worked on the very first try, a clean and believable integration.

Matching the light, and walking into the forest

The same lighting fix applied to another forest frame, with roles assigned to each reference (character / background).

We used the same approach for the forest trail. We combined the two references again, with a prompt in which she walks with her back to the viewer, on a 24mm lens, background slightly blurred, morning light, and again asked the model to match the figure's lighting to the scene. She walks away from camera, deeper into the forest. This rear view shifts the focus from her face to the journey and sets the direction of movement for the climb that follows.

Matching the light, and walking into the forest

Rear view on the trail, 24mm: face hidden, direction of movement set toward the climb.

Tip: When you place a character into a scene, always ask the model to adjust her lighting and color temperature to the location. Otherwise she looks pasted on. This one line is what makes her sit in the frame.

 

When it does not work: iterate, or switch the model

Not every frame comes easily. For the transition from the forest toward the mountains we wanted the heroine to push branches aside and come through, an extreme close up on a 24mm lens, low angle, looking up. The model just would not understand what we wanted. We got there through a series of small prompt changes and repeated generations. For this single frame we did 17 generations before we got the result.

When it does not work: iterate, or switch the model

The branches shot that took 17 generations: extreme close-up, low angle, looking up.

By contrast, the mountain shots that followed came together in just a few takes. With all the earlier trial and error behind us, the work went noticeably faster. So the difficulty is not predictable, and that is simply part of working with these tools.

Tip: Do not abandon a frame if it does not come out on the first or second try, and do not expect one prompt to give a perfect result instantly. Tweak the prompt in small steps and keep generating. And if one model keeps stumbling, switch to another version just to see if it handles it better.

 

The hero shots

The whole video builds toward the product, so for the climax we made two hero frames of the bottle.

For the first we placed the bottle on a flat stone at the summit at sunset. We took the product image of the perfume bottle and asked the model to set it into the mountain location with one strict instruction: keep the exact shape, proportions, and label of the bottle. The low sun sits directly behind it, so the light passes through the glass and the liquid and the bottle glows. A classic backlit hero shot. After the cold dawn of the opening, this warm golden-hour finish is the emotional release, the moment the fragrance opens.

The hero shots

Hero shot one: the bottle backlit on a summit stone at sunset, its shape and proportions preserved.

For the second we wanted the heroine holding the perfume. The main difficulty here is that the model does not know the real size of the bottle relative to the hands, so it made the bottle too big or too small and the grip looked unnatural. Our solution was to find a Pinterest reference of hands holding this exact perfume, attach it, and ask the model to repeat the hand position. That fixed the scale and the grip. In the frame we liked, the model had also added a ring on her finger that we did not want, so we took the finished frame and wrote a simple prompt: keep everything as is, only remove the ring. Another precise local edit instead of a full regeneration.

The hero shots

Building the in-hands hero: photo 1 a reference of hands holding a similar bottle (for scale and grip), photo 2 the character, photo 3 the bottle.

The hero shots

Hero shot two: the bottle in her hands, scale fixed with a hand reference, the stray ring removed with a local edit.

Tip: For an object held in the hands, give the model a reference of the hand position, because it cannot judge scale on its own. And when a frame is otherwise perfect but the model slipped in an artifact, fix that one thing locally instead of regenerating everything.

 

A quick reference: camera keywords we use in prompts

These are the shooting parameters we lean on most when building a frame: angle, shot size, lens, and light. Adding them to a prompt is what turns “a girl in a forest” into a directed shot. Below is our working shortlist.

Angle and orientation: low angle (subject looks dominant, figure lengthened), high angle (opens the ground, adds vulnerability), eye level (natural proportions, direct contact), rear view (face hidden, good for walking into a scene and setting direction), three-quarter view (the classic flattering portrait angle).

Shot size: extreme wide (tiny figure, epic scale), wide (full figure, establishes place), medium (waist up, body language), close up (face fills the frame, emotion), extreme close up (a fragment like an eye or lips, intimacy or tension).

Optics: 24mm wide angle (wide scene, depth, used on the forest trail), 50mm (closest to the human eye), 85mm (portrait standard, soft background separation), shallow depth of field or f1.8 (blurred background, subject isolated).

Light: golden hour (warm soft light, our whole finale runs on it), backlighting (source behind the subject, glow, the bottle hero shot), rim lighting (thin edge light separating subject from a dark background), crepuscular rays (visible sun rays through clouds), cinematic lighting (deep shadows, film feel).

 

Lay the frames out before you animate

Once we had basically finished generating the still frames, we opened a video editor and laid them out roughly in the order we would animate them, just to see the whole thing at a glance and check whether any shots were missing. We saw we wanted a couple more: a frame of the heroine stepping toward the mountains, and one at the very end, her carrying the fragrance away in her backpack, which felt like a great closing shot. So we generated those afterward.

Lay the frames out before you animate

Laying all the frames out in the editor before animating, to see the whole piece and spot the gaps.

Tip: Yes, you can plan everything in your head, but it is still worth doing this layout so you can see the whole piece and spot what is missing. Some of the best ideas come during the work itself. You cannot always plan everything up front.

 

Reference: image prompt keywords (Midjourney/Nano Banana)

The keyword shortlists we keep for the image stage, for building a frame from the character map: angle, optics, light, and shot size, with what each one does.

Camera angles and perspective

Keyword

Effect

low angle shot

Camera below eye line, aimed up. Lengthens vertical lines, gives dominant status.

high angle shot

Camera above subject, aimed down. Reduces significance, emphasizes vulnerability.

eye-level shot

Neutral position, maximum portrait likeness and natural proportions, direct-contact feel.

overhead shot

Camera perpendicular to the ground (90 deg). Ideal for diagrams, flat lays, geometric comps.

bird's-eye view

Broad panorama from high up; emphasis shifts from character to landscape and architecture.

worm's-eye view

Extremely low angle, near ground level. Exaggerates scale, making objects giant.

dutch angle

Tilted horizon.

over-the-shoulder shot

Classic dialogue angle; blurred foreground silhouette (shoulder, head) focusing on the other person.

pov (point of view)

Imitation of seeing through the character's eyes.

drone shot

Dynamic aerial angle combining height and wide angle.

profile shot

Strictly from the side. Emphasizes silhouette, contours of nose and chin.

three-quarter shot

Subject turned 45 deg to camera. Standard classic portrait, adds volume to the face.

rear view

Camera behind the subject, face hidden. Emphasis on context, clothing, or mystery.

ground-level shot

Camera on the surface, foreground (grass, asphalt) heavily blurred. Creates depth.

satellite view

Extreme height, cartographic style. Detail dissolves into landscape textures.

hip-level shot

Angle characteristic of westerns.

shoulder-level shot

Slightly below eye level.

god's eye view

Strictly vertical view down onto a vast scene.

wide-angle perspective

Pronounced perspective distortion, stretched frame edges, enlarged foreground objects.

fisheye view

Strong spherical distortion (180 deg). Action-cam / 90s-clip styling, edge vignetting.

macro perspective

View into the micro-world; small details (insects, drops) become giant, background blurred.

selfie angle

Characteristic close-range face distortion, outstretched arm often visible.

security camera view

High angle, grain, specific color grading, slight edge distortion.

Optics, lenses and film

Keyword

Effect

35mm film

Classic film aesthetic: light grain, natural color, high dynamic range.

imax

Ultra-high resolution, epic scale, deep detail across the frame, near-square format.

polaroid

Vintage style, soft focus, characteristic chemical color, often a white border.

drone photography

High sharpness, wide angle, emphasis on geometric landscape patterns.

mirrorless

Like DSLR, but often read as a more clinically clean digital image.

telephoto lens

Strong perspective compression, blurred background, subject isolation. Ideal for portraits or wildlife.

wide-angle lens

Captures a wide scene, stretches edges, enhances perspective and spatial depth.

macro lens

Extreme detail of small objects, very shallow depth of field.

fisheye lens

180 deg spherical distortion, vignetting, artistic deformation.

tilt-shift lens

Miniature effect (toy town), sharp focus band, strong blur top and bottom.

50mm lens

Classic lens; perspective closest to the human eye.

85mm lens

Gold standard of portraiture; soft background separation, correct face proportions.

f/1.8 aperture

Shallow depth of field, strong background blur, emphasis on the eyes, soft look.

f/16 aperture

Large depth of field; all planes in focus (foreground to infinity). For landscapes.

bokeh

Aesthetic quality of blurred out-of-focus light sources (circles, highlights).

motion blur

Smearing of moving objects against a static background, or the reverse. Dynamism.

double exposure

Artistic overlay of two images (usually silhouette plus landscape).

long exposure

Light trails from headlights; a sense of flowing time in a static frame.

kodak portra 400

Warm tones, excellent skin-tone rendering, fine pleasant grain.

fujifilm velvia

High color saturation, high contrast, cool tones. Ideal for nature.

ilford hp5

Classic black-and-white film; high contrast, deep shadows, pronounced artistic grain.

vhs tape

Glitch effects, color noise, scan lines, low resolution. 80s to 90s aesthetic.

cctv footage

Low quality, compression, characteristic angle, timestamp, cool tones.

thermal imaging

Image in the infrared spectrum.

night vision

Green monochrome, heavy grain, glowing eyes, vignetting.

anamorphic lens

Cinematic widescreen look, oval bokeh, horizontal lens flares from light sources.

Lighting

Keyword

Effect

golden hour

Warm, soft, golden light (first hour after sunrise, last before sunset).

blue hour

Cool, deep blue twilight. Low contrast, melancholic mood.

cinematic lighting

Accented scheme with deep shadows and high detail; imitates film production.

volumetric lighting

God rays: light made visible by particles in the air (fog, dust).

studio lighting

Controlled, perfect light; often implies several sources (key, fill, back).

neon lighting

Bright, saturated artificial colors (cyberpunk). Colored reflections on skin and surfaces.

natural light

Light from sun or window. Soft, realistic, balanced in color temperature.

backlighting

Light source behind the subject. Creates silhouette or halo (rim light), hides facial detail.

rim lighting

Narrow strip of light along the subject's edge, separating it from a dark background.

hard lighting

Direct light (midday sun, flash). Sharp deep shadows, high contrast, drama.

bioluminescence

Natural cold glow of living organisms (fungi, plankton) in darkness. Mystical.

candlelight

Warm, flickering light. Creates intimacy or mystery, deep shadows.

ambient light

Even lighting with no clear directional source. Soft shadows, low contrast.

moody lighting

Mostly darkened scene with rare accent pools of light.

lens flare

Bright geometric artifacts from light hitting the lens directly. Adds optical realism.

silhouette

Subject fully black against bright light. Details hidden; only shape matters.

split lighting

Face lit exactly in half, other half in deep shadow. Dramatic.

butterfly lighting

Light source above and directly in front of the face.

diffused light

Light through clouds or white fabric. No harsh shadows, softness.

spotlight

Narrow directed beam highlighting one detail in the dark. Theatrical.

overcast lighting

Even, shadowless light from the whole sky. Cool tones, softness, no contrast.

crepuscular rays

Visible sun rays breaking through gaps in clouds. Epic atmospheric effect.

strobe light

Club flash effect. Frozen movement, high sharpness, cool tint.

Shot size, composition and framing

Keyword

Effect

extreme close-up

Focus on part of a face (eye, lips) or object. Maximum macro texture detail.

close-up

Face and neck fill the frame. Emphasis on expression and emotion, background heavily blurred.

medium close-up

From the chest up. Classic portrait format, part of shoulders visible.

medium shot

From the waist up. Shows interaction of hands and body, body language.

cowboy shot

From mid-thigh (historically to show the pistol). Adds confidence and dynamism.

full body shot

Whole figure with legs and shoes. Important for showing an outfit.

wide shot

Character occupies a small part of the frame, lots of surroundings. Scene context.

extreme wide shot

Character tiny; emphasis on epic landscape and the scale of the world.

long shot

Like wide, but the emphasis is on the distance from camera to subject.

establishing shot

General view of the location (city, house) before a scene. Sets atmosphere.

two-shot

Two characters in frame. For dialogue and relationship interaction.

group shot

A group of people (three or more).

crowd shot

Many people.

selfie

Shot at arm's length; characteristic face distortion, outstretched arm often visible.

candid shot

Unposed photo, naturalness, a caught moment of life.

action shot

Frozen movement, dynamism, background blur, muscle tension.

hero shot

Usually low angle, epic pose, advertising style, idealization of the hero.

product shot

Emphasis on the object, clean background, perfect studio light.

flat lay

Strictly top-down view of objects artfully arranged on a surface.

symmetrical composition

Object strictly centered, perfect left/right balance.

rule of thirds

Subject offset from center to the third-lines. Harmonious composition.

centered composition

Object exactly in the center of the frame.

framing

Shooting through a natural frame (window, arch, branches) that creates depth.

negative space

Lots of empty space around the subject. Minimalism, sense of air and solitude.

shallow depth of field

Background heavily blurred (bokeh), subject visually separated (f/1.8).

point of view

View through the eyes of a participant, hands in frame.

 

Reference: angles, shots and camera height

A shorter companion list: subject orientation, shot size and camera angle for the image stage.

Subject orientation

Keyword

Effect

front view

Subject faces the camera directly. Open, confrontational, symmetrical; full face visible.

side view / profile

Subject seen from the side. Emphasizes silhouette, nose and chin contour, gaze off-frame.

rear view

Camera behind the subject, face hidden. Emphasis on context, clothing, mystery; good for walking into the scene.

three-quarter view

Subject turned 45 deg to camera. Adds volume and depth to the face; the classic flattering portrait angle.

over-the-shoulder view

Framed past the subject's shoulder. Blurred foreground shoulder or head, focus beyond; ties viewer to their point of view.

top view / high angle

Camera above, looking down. Diminishes the subject, opens up the ground plane and surroundings.

low view / from below

Camera below, looking up. Makes the subject monumental and dominant, lengthens the figure.

first-person / POV

The scene as seen through the character's own eyes; hands or gear may enter the frame. Immersive.

inside view

Camera placed inside an object or enclosed space looking out (tent, hood, car, foliage). Frames the world through an opening.

Shot size

Keyword

Effect

extreme wide shot

Subject tiny in a vast environment. Emphasis on epic landscape and the scale of the world.

wide / long shot

Full figure with lots of surroundings. Establishes place and the subject's relation to it.

medium shot

From the waist up. Shows body language and the interplay of hands and body; conversational distance.

close-up

Face and neck fill the frame. Emphasis on expression and emotion, background heavily blurred.

extreme close-up

Focus on a fragment (eye, lips, texture of an object). Maximum macro detail, high intimacy or tension.

Camera angle

Keyword

Effect

high angle

Camera above eye line, aimed down. Reduces the subject's significance, suggests vulnerability, opens the ground.

low angle

Camera below eye line, aimed up. Grants dominance and power, lengthens vertical lines.

dutch angle

Tilted horizon. Adds unease, tension or dynamism; a deliberately off feeling.

eye level

Neutral, natural height. Maximum likeness and honest proportions, direct eye contact with the viewer.

 

Step 3. Animation

As our main animation tool we chose Kling 3.0, because it gave us the best balance of price and quality.

 

Create Elements first

Before animating anything, one important preparation step: create Elements. An Element is a saved, reusable reference. It works on the same principle as the character sheet, but for the video stage. We created two: a Character element built from the 360 sheets we had already made, and a Prop element for the bottle. Then, when animating a shot, we attach these Elements and reference them right in the prompt with the @ symbol, and the model keeps the character and the product consistent across the whole video.

You build an Element by feeding it your references. For Redhead-girl we used the full body turnaround, the portrait grid, and a clean close up. That is the direct payoff of the earlier work: the character sheet from the image stage becomes the foundation of the video Element. The cleaner your reference, the more stable the character stays in motion.

Create Elements first

Building the @Redhead-girl Element from the 360 sheets made at the image stage.

Tip: This completes a consistency system that runs through the whole pipeline: the 360 sheet at the image stage, lighting-matched integration into scenes, and now Elements plus start-and-end frames at the animation stage. The heroine and the bottle never drift. For an ad, locking the product is just as important as locking the character.

 

Start and end frames

We started with Kling 3.0, the newest video model at the time. For the very first shot, the heroine opening her eyes, we used a start frame (eyes closed) and an end frame (eyes open), so the model would animate the movement between them while keeping her exact face. The prompt describes only the motion, not the whole scene: the woman slowly wakes, a soft single blink, a quiet waking breath, a few red curls shift, a calm unhurried awakening, camera almost still with an extremely slow subtle push in. One catch: when you work with start and end frames, Kling does not add music, so you handle sound separately in the edit.

Start and end frames

The waking shot in Kling 3.0 using a start frame (eyes closed) and an end frame (eyes open).

Tip: To preserve your specific character in motion, give the model both the first and the last frame. And write animation prompts as movement, not description: action, camera move, tempo, and limits, since the model already sees the picture.

 

The drone shot, and finding the story in the edit

For the establishing shot we wanted a drone move over the tent, and it came together on the second try. The prompt was an aerial smooth forward flight: the camera glides over the tent and passes above it, the tent drops into the lower part of the frame as a misty autumn forest opens up ahead, warm lantern glow from the entrance below, fog drifting between the trees, cold pre-dawn light, steady, no flicker.

The order shifted as we worked. At first we wanted to open on the heroine's eyes, but we decided to start with this wide drone shot instead, because it felt like a stronger opening. We added a whoosh sound as the camera flies in, so that right after it her eyes snap open, as if something woke her, as if something called her to the mountains. That idea came to us during the animation itself.

Tip: Name your camera move precisely (aerial, tracking, dolly, orbit) and describe how the composition changes through the shot. Add “steady, no flicker” when even light matters. And do not be afraid to reorder shots once you see them live, or to use sound as an accent on a cut, so the montage lands and even creates meaning rather than just playing underneath.

The drone shot, and finding the story in the edit

The drone establishing shot in Kling 3.0: an aerial fast forward flyover over the tent, fog and leaves streaking past.

 

Take the best moment

For the boots shot we had written that she ties a bow on her laces, and of course the laces started morphing and doubling. Anatomically correct shoelace tying is very hard for a neural network. But it was not critical, because we only needed about a second from that clip. The model added some junk, but the one second where she pulls the bow tight came out clean, and that second was all we needed.

Tip: You do not need a perfect clip, you need a perfect moment. Pull the best one or two seconds and move on. And expect fine hand mechanics (tying laces, fingers, precise manipulation) to be hard, so plan a short usable fragment where the artifact does not show.

Take the best moment

The boots shot in Kling 3.0: hands tying the bootlace. Only the one clean second where the bow pulls tight went into the cut.

 

Simple actions are easy

The next shots came exactly as we wanted, quickly and without trouble, because a neural network animates simpler actions really well. Two example prompts:

“She grips the backpack strap and settles it on her shoulder, subtle tug, red curls stir, slow push in. Unhurried. No warping hands, no morphing straps.”

“She walks away deeper into the forest, tracking shot following her, slow dolly in, fog drifts, leaves fall. Unhurried, luxury pace. No warping, no shake.”

Tip: Build the prompt as action + camera move + tempo + limits, add small details like hair, fog, or leaves for life, and write in your restrictions ahead of time (no warping hands, no morphing straps, no shake). Preventing artifacts is cheaper than regenerating.

Simple actions are easy

A simple action in Kling 3.0: she walks away deeper into the forest, tracking shot, slow dolly in.

 

The climb: switch model, and budget for it

We added a shot of the heroine climbing a rock, which was not in our original storyboard, but we felt a moment of effort just before the finale would work on contrast: the struggle, and then the sunset and the fragrance as a reward. Kling could not animate it, the climb came out unnatural. So we tried Seedance, and even though it is about three times more expensive per token, it animated the shot well on the first try and the climb looked natural.

Tip: Use the contrast of effort then reward before a finale. And budget for it, because some complex movements are hard for Kling and you should set aside money for a model that costs roughly three times more.

The climb: switch model, and budget for it

The climb, animated in Seedance 2.0 after Kling could not manage it: reaching for the next rock, heavy breathing, a light handheld feel.

 

Kling 2.6 for slow motion

Some shots we animated in Kling 2.6 rather than 3.0. Kling 3.0 is tuned for movement and is not great at slow, gentle motion, like hair drifting slowly. When we tried slow drifting hair in Kling 3.0 it came out like a hurricane rather than a light breeze, which obviously did not work. So we moved those shots to Kling 2.6, which handles slow gentle movement far better. It has a downside though: it tends to add strange things of its own. On the shot where the heroine sprays the perfume it kept adding weird artifacts, some kind of smoke or exhale, so we ran several generations before we got what we wanted.

Tip: Use Kling 2.6 for slow motion and very smooth shots, but only ones that are simple in movement, since it invents artifacts on complex actions. And pick not just a model but a version to match the character of the movement: active, gentle, or complex, each has its strength.

Kling 2.6 for slow motion

Slow drifting hair in Kling 2.6, which handles gentle motion far better than 3.0.

 

Animating the hero shots of the flacon

The first hero shot is the bottle standing on the stone at sunset. For the animation we did not use any reference. We just described what we wanted in words: the low sunset light behind it, and where we wanted the glint to sit on the glass. It came together fast. We picked a take we liked from the very first ones, and that is the one that went into the final cut.

Animating the hero shots of the flacon

Hero shot one: the bottle on the summit stone, the frame we then animated with “Turn to video”.

A near-static product hero like this often comes quickly, because there is almost no movement that can break. A clear description of the light does most of the work.

Tip: For a mostly static product hero you do not always need a reference. Describe the light and the exact highlight you want in words. These shots tend to land in the first few takes, so do not overbuild them.

The second hero shot is the bottle held in her hands, and this one gave us trouble. In Kling 3.0 the animation kept spinning the bottle around in her hands far too much. As we mentioned earlier, Kling 3.0 is built for strong movement, so it overdid a delicate action and it looked strange. We moved to Kling 2.6, which is gentler, and prompted it plainly: hands slowly turning the bottle, a glint travelling across the glass, the camera doing a half-orbit, a faint smile, and, as always, no music, natural sounds, and she does not speak. We ran a few generations and pulled the best couple of seconds out of one of them.

Animating the hero shots of the flacon

Hero shot two: the bottle in her hands, animated in Kling 2.6. This one was cut from the final.

In the end, though, this shot did not make the final cut. We realised we had two similar medium shots of the perfume, the bottle on the stone and the bottle in her hands, and running them one after another felt flat. So we built the finale on variety instead: the bottle on the stone, then she sprays it, then she carries it away in her backpack. The in-hands clip came out fine, it just did not earn its place in that sequence.

Two lessons sit in this one shot. A delicate, slow action belongs in Kling 2.6, not the movement-heavy 3.0. And a clip being good is not enough on its own: two shots of the same size and subject back to back kill the rhythm, so it has to serve the sequence or it goes.

Tip: Match the model to the motion, gentle actions to Kling 2.6 rather than 3.0. And be ready to drop a clip you like: if it doubles another shot's size and subject, the film is better without it.

 

Reference: animation prompt keywords (Kling/i2v)

A keyword table for the animation stage, for bringing a finished frame to life. It pairs with the shooting-keyword shortlists from the image stage above.

Camera movement

Keyword

Effect

pan left / pan right

Camera rotates left or right while staying in place. Good for surveying a space.

tilt up / tilt down

Camera looks up or down. Conveys height, scale, motion from bottom to top.

dolly in / push in

Camera physically moves toward the subject. Heightens emotion and detail.

dolly out / pull out

Camera moves back. Reveals the surroundings and scene context.

tracking shot

Camera moves parallel to a moving subject. Good for walk-alongs, running, profile movement.

truck left / truck right

Camera moves left or right parallel to the scene. Creates parallax.

orbit / arc shot

Camera rotates around the subject. Good for presenting a product or character.

crane up / pedestal up

Camera rises vertically. Good for opening and large-scale shots.

roll

Camera rotates around its own axis. Adds disorientation or dynamism.

handheld camera

Slight shake, live-operator feel. More documentary and realistic.

shaky cam

Strong shake. For action, running, chaos. Use sparingly.

static shot / tripod

Camera is still. Use when you want movement within the frame, not camera movement.

Zoom and focus

Keyword

Effect

zoom in

Optical zoom in; focal length changes, background visually compresses.

zoom out

Optical zoom out; field of view widens.

rack focus

Focus shifts smoothly from foreground to background or the reverse.

shallow depth of field

Blurred background, sharp subject. Cinematic.

deep focus

Everything sharp, foreground and background. Good for landscapes and complex scenes.

bokeh

Beautiful blurred circles of light in the background.

dolly zoom / vertigo

Camera pull-back combined with a zoom in. Complex but expressive.

Speed and motion

Keyword

Effect

slow motion

Slowed movement. Must be specified up front, at generation.

hyperlapse / timelapse

Accelerated time: clouds race, light shifts, flowers bloom.

speed ramp

Speeding up or slowing down within one shot. A stylistic device.

frozen time / bullet time

Camera moves around a frozen subject, the “Matrix” effect.

fluid motion

Smooths the movement of liquid, fabric or hair; removes jerkiness.

dynamic action

Pushes the model into sharper, sweeping movements.

fast pace

Increases the overall speed of what is happening in the frame.

 

Step 4. Assembling the film in CapCut

By this stage we had a folder of separate animated clips, and the edit is where they finally turn into a film. We want to open this chapter with the idea that shaped the way we think about montage: the Kuleshov effect.

It comes from an experiment in early Soviet cinema. Kuleshov took a single neutral shot of an actor's face and cut it next to three different images: a bowl of soup, a child in a coffin, a young woman lying down. Each time the audience read a different emotion off the same face, hunger, then grief, then desire, and admired how expressive the actor was. But it was the identical shot every time. The feeling was never in the face. It came from whatever sat next to it.

Assembling the film in CapCut

The Kuleshov experiment: one neutral face intercut with soup, a coffin and a reclining woman, read as hunger, grief and desire.

That is exactly why montage matters. The emotion is not locked inside one frame, it is created by how the frames sit against each other, so arranging them in the right order is not a finishing touch, it is where the feeling is made.

We used this directly. Our heroine looks off at something, and right after her we cut to the flacon on the summit. Because of that order, her expression reads as one clear emotion: this is her reward, the thing she climbed all that way for. On its own her look is neutral. It only becomes "she found what she was searching for" because the bottle lands right after it.

Our timeline in CapCut: the heroine's reaction cut straight into the flacon.

Tip: The emotion is not only inside a frame, it is created between frames. Before you cut, think about what each shot sits next to. The order is where the meaning comes from.

 

The montage in CapCut

For the montage we used CapCut. It had everything we needed for a short vertical reel: speed curves, cutting on the audio waveform, simple colour balancing, and a clean 9:16 export.

 

Rhythm, and speed curves

We laid the clips out in story order and cut each one down, keeping only the best one or two seconds from each 5 second generation. From there we wanted the film to feel more dynamic, so we worked on the speed. Some clips we simply sped up, mostly the connective moments like walking. On others we added a speed curve, which ramps the speed up and down inside the shot instead of running it at one flat rate.

That tool is worth knowing. A small jump in speed inside a shot catches the eye and keeps the viewer's attention, which is exactly what a short piece needs to hold. It is also the honest way to shape tempo, because AI video cannot add convincing slow motion after the fact, so you build the pace in the edit instead.

Tip: To make a short film feel dynamic, do not leave every clip at one flat speed. Speed up the filler, and use speed curves to ramp the pace inside a shot. That change of speed is what holds the eye.

Rhythm, and speed curves

The speed-curve panel in CapCut, with a ramp applied to a clip.

 

Choosing the music

Once our animated clips were laid out in the right order, and we had checked that the shot sizes alternated the way we wanted, we moved on to choosing the music. Our first thought was to generate a track. We tried it in ElevenLabs and it actually came out pretty good. But the finished download was paid, and we had already spent a lot on the video itself, so we decided not to. Instead we went to Pixabay and downloaded free music there.

Choosing the music

Pixabay: free, commercially usable music. Watch the licence.

For a spec piece the soundtrack does not have to be generated or paid for. There is clean, free, commercially usable music out there, so there is no reason to blow the rest of the budget on the track.

Tip: Sort the music out early, and check you are actually allowed to use it commercially. Free libraries like Pixabay cover a spec piece at no cost, but watch the license: CC-BY-NC means no commercial use.

 

Cutting to the music

We cut the clips to the music so the tension would build. The track starts slower and keeps rising through the film. At the exact moment the heroine sees the perfume there is a sharp whoosh, and the music drops away almost to nothing. Then, at the very end, when she carries the fragrance off in her backpack, the music comes back in, laid over the shot as a separate track.

Cutting to the music

The finished timeline: the music under every clip, a separate audio clip at the end, and the endcard on the final backpack shot.

The sound is doing the emotional shaping here, not just filling the background. Building the music up to the moment she finds the bottle, cutting it to a hush on the whoosh, and letting it swell back on the ending is what makes the climax and the release actually land.

Tip: Cut your clips to the music, not the other way around, and let the track's dynamics carry the story: build toward the key moment, drop to near silence on the accent, then bring the music back for the resolution.

 

Sound: what the model gave us, and what we added

The mechanism here is easy to get wrong, so it is worth spelling out. Kling only gives you a silent clip when you generate a shot from both a start frame and an end frame. That is exactly why we used a start-and-end frame for just one shot, the one where her eyes have to open. Every other shot we generated from a start frame only, with the actions she performs written into the prompt, and those clips come out with sound. So her footsteps as she walks, the spray of the perfume, all the natural sounds, came straight out of the generation, and we kept them instead of rebuilding them in CapCut.

On those start-frame prompts we always wrote in two things: no music, and the natural sounds we wanted. That way Kling gives us clean foley without laying its own music track under the shot, because the music is ours to add in the edit. We also learned to write that she does not speak, unless we actually want a voice, otherwise the model can invent one.

What we added in the edit was only the music. Everything else came straight out of the generation: the footsteps, the spray of the perfume, and the whole ambient bed of the film, the damp forest, the wind, the birdsong. Kling produced all of it on the start-frame shots, and we just let it play under the music.

Sound carries emotion as much as the image does, and here the model did most of that work for us. There was no point rebuilding a soundscape it already generated. Our job in the edit was the music, and letting it sit against the sound the clips already carried.

Tip: Only start-and-end-frame shots come out silent in Kling, so use that pairing only where you truly need it (like eyes opening) and generate everything else from a start frame, so the natural sound comes for free. In those prompts write in "no music" to get clean sound without a music bed, and add "she does not speak" when you do not want the model to invent a voice. Then all you add in the edit is the music.

 

The whoosh into the waking: the cut that makes meaning

We reordered the opening. Our plan was to open on the heroine's eyes. In the edit we decided the wide drone shot over the tent was the stronger opening, so we led with it and put a whoosh on the camera move. Right after the whoosh, her eyes snap open, as if something called her, as if the mountains woke her. That idea only arrived while we were editing.

This is the Kuleshov effect working for us directly. The meaning is not inside either shot. It appears when we put them next to each other, with the whoosh doing the work on the seam. Our summit reaction shot behaves the same way: her face reads as "she found the reward" only because it sits after the climb and before the flacon. That is exactly why the finale order (reaction, then the bottle) mattered so much.

Tip: Do not be afraid to reorder shots once you see them moving. Put a sound accent on the cut itself, so the seam creates meaning instead of just passing. Some of the best ideas come during the edit, not the plan.

 

Matching color across the shots

You can also color-correct the clips right in CapCut, and we needed to. Even though we generated everything in one palette, the shots still did not match perfectly out of the model. For example, the forest shot where the heroine pushes the branches aside came out a little warmer than the rest, warmer than we wanted, so we brought the temperature down a touch on that one clip. It works like Photoshop, temperature, light, and the rest, except you are adjusting video instead of a still. We also pushed the whole film along the light arc: cool at the start, warming toward the golden summit.

Matching color across the shots

CapCut's colour-correction panel, cooling the temperature on the forest shot.

A single shot that falls out of the color instantly reads as a cut and breaks the spell. Being able to correct each clip is what pulls everything into one consistent tone, which is exactly what you want across a whole film.

Tip: Do not count on one palette at generation to give you a matched film. Color-correct in the edit, clip by clip, the way you would retouch a photo, until every shot sits in the same tone and you cannot see the joins.

 

Not everything good makes the final cut

We want to add a thought here. You do not have to put everything you generated into the film, not even the shots that came out well. Sometimes a good clip has to go, precisely so the piece stays dynamic and the story stays sharp.

Tip: Judge each clip by what it does for the whole film, not by how nice it looks on its own. Letting go of good material is part of editing, and it is what keeps a short piece tight.

 

Straight cuts

We did not use any special transitions between the shots. For the calm, editorial goal of this piece, plain straight cuts were exactly right.

 

The closing shot

For the very end we added a frame that was not in our first storyboard: the heroine walking away with the flacon tucked into her backpack pocket, carrying the fragrance back down with her. It came to us while we were laying the film out, and it felt like the right closing beat, the journey completing.

A closing image that resolves the arc, she found it and now she carries it home, gives the film an ending instead of just a stop.

Tip: Leave room for the ending to reveal itself in the edit. The closing shot that resolves the story is often not the one you planned.

 

Do and Don't: the whole pipeline at a glance

Do

  • Build a real moodboard: collect the exact shots and angles you need, and your storyboard almost writes itself.
  • Lock consistency early: a 360 character sheet covering the face and the full figure, clothing and backpack, plus Elements for the character and the product, so nothing drifts.
  • When you drop a character into a scene, always ask the model to match her lighting and color temperature to the location.
  • Attach both the character and the location reference, and check both actually loaded before you generate.
  • Build prompts as action plus camera move plus tempo plus limits, and name the camera move precisely (dolly in, tracking, orbit, tilt up).
  • Write your restrictions in ahead of time (no warping hands, no morphing straps, no shake). Preventing artifacts is cheaper than regenerating.
  • Use both a start and an end frame when you need to preserve a specific face in motion, like eyes opening.
  • On start-frame shots, write "no music" for clean natural foley, and "she does not speak" if you do not want a generated voice.
  • Match the model to the motion: Kling 3.0 for strong movement, Kling 2.6 for slow and gentle, Seedance for the hard shots.
  • Lay all your frames out before animating, to see the whole piece and spot what is missing.
  • Cut to the music and let its dynamics carry the story: build to the key moment, hush on the accent, swell back for the resolution.
  • Color-correct clip by clip in the edit until every shot sits in one tone.
  • Be ready to cut good material if it does not serve the rhythm or the sequence.

Don't

  • Don't treat the moodboard as a collage of pretty pictures; it is a planning tool.
  • Don't fix only the face on your character sheet, or the model will invent new outfits in later scenes.
  • Don't rely on the model to remember the atmosphere on its own; without the location reference the light drifts.
  • Don't expect a perfect frame in one shot. Build the composition, then swap in your consistent character and relight.
  • Don't abandon a frame after one or two tries. Tweak in small steps, and switch model versions if one keeps stumbling.
  • Don't animate a delicate, slow action in Kling 3.0, it over-moves it; and don't ask Kling 2.6 for complex motion, it invents artifacts.
  • Don't chase a perfect clip. Take the best one or two seconds and move on, especially for fine hand mechanics.
  • Don't use a track you cannot license commercially (watch for CC-BY-NC).
  • Don't rebuild sound the model already generated, like footsteps or the spray.

 

What this whole project taught us

If we gather it all into one thought, it is this: the film is not generated, it is assembled. The AI gave us the individual frames, but the story, the emotion and the rhythm all happened in the edit. Two shots placed next to each other can carry a feeling that is in neither of them on its own.

Everything else follows from that. Match the tool to the task, and even to the motion: Midjourney for art, faces and locations, Nano Banana for edits and compositing, Kling 3.0 for strong movement, Kling 2.6 for gentle motion, Seedance for the hard shots. Lock your character and your product early with the 360 sheet and the Elements, so they never drift. Keep one palette and pull every clip into a single tone. Build the sound and the music on the same arc as the light. And be ready to cut good material, because a short film lives or dies on its rhythm.

The tools will keep changing every few months. The method is what stays: think about the edit while you generate, and treat the cut as the place where you actually direct the film.

The AI generated our frames. The film was made in the cut.


Anfisa Kosenkova
UX/UI Designer
image
Expertise
Question to the expert
image

We have available resources to start working on your project within 5 business days

1 UX Designer

image

1 Admin

image

2 QA engineers

image

1 Consultant

image