Video Prompts
Switch the generator to Video, pick a source (one of your images, or an e621 / e6ai post link) and choose how much you want to write. The image already exists, so a good video prompt describes motion: what moves, how it moves, the light and the sound.
There are two modes. Auto writes the whole prompt from the picture's tags, so you type nothing. Manual sends your text through untouched, which is the one to reach for when you know exactly what should happen.
Available on video-enabled plans
Length, and what it costs
A clip is 5 or 10 seconds. One credit buys 5 seconds, two buy 10. Ten seconds is not room for a second action: it is the same action given a slower, fuller arc, or one repetitive cycle running longer. Asking for two things in ten seconds is the fastest way to get a cut, and a cut is where clothing and anatomy reset.
You can also set an optional end image from your gallery, and the clip will land on it. And once a clip is done, Continue this clip starts a new one from its last frame, so a sequence stays coherent. A continuation is always written by hand: its first frame comes from a video, which carries no tags for Auto to read.
The One Rule
Describe what moves, not what things look like. The image already shows the character, colors, outfit and scene, so re-describing them just pushes the model to redraw them, which causes morphing. Say what happens next, keep it to one small action, and keep the camera still.
What to Describe
| Part | What to write | Example |
|---|---|---|
| Action | The single thing that happens | she slowly turns to the camera |
| Body response | How her body reacts to it | her ears flick, her tail sways |
| Pace | Natural and a little slow (never fast or sudden) | slow, steady, unhurried |
| Lighting | One short clause: source and warmth | soft warm window light from the left |
| Sound | What you hear | soft breathing, a quiet room |
| Speech (optional) | A short line with an emotion label | [breathlessly] "come here" |
Example
One character, she slowly turns to the camera and smiles, her ears flicking and tail swaying gently, soft warm window light from the left, quiet breathing and a soft room tone, static camera, single continuous shot, clothing unchanged, same art style, no extra characters or limbs, no background music, no censorship, stable face, no distortion, no jitter, clean loop.
The tail from static camera onwards is the quality clause. Auto appends it for you; in Manual you have to paste it at the end of your own line. That is not optional politeness: the model has no separate negative-prompt channel, so this clause is the only thing holding anatomy, coverage and style still. Each part of it shuts down one specific failure, listed below.
Keep the camera still
A completely static camera is the single highest-success choice. Avoid pans and especially orbiting shots, which fail most of the time on short clips. Only the subject should move. The one exception is a scene that is already about movement (running, falling, swimming), where a single slow push in or out can work.
One shot, no cut
Left to itself the model will happily cut to a second angle inside your clip, and a cut is where things break: clothing comes back, anatomy changes, the character stops looking like herself. Ask for a single continuous shot and never describe a second angle or a "then it switches to".
Say how many characters
Open with the count in plain words (one character, two characters) and it stops the model inventing an extra hand, limb or body part to fill the frame. It is the single most effective fix for unwanted anatomy appearing mid-clip. Do not reuse the tag words solo or duo here: they read as style tags rather than a count, and on a two-character scene that can cost you a character.
Don't let the style drift
The clip should keep the exact look of your source image. Do not name an art style in the prompt, not even the one your image already has: naming one is how a 2D piece comes back rendered as something else. same art style in the tail handles it.
Lighting matters most
Naming the light is the biggest quality lever. One short clause is enough: the source and direction plus a warmth, for example soft warm window light from the left, low candlelight, deep shadows, or cool blue neon. Leaving the light unspecified is the most common cause of a flat, muddy result.
Sound and speech
Your video has audio. Add one short clause of sound that fits the action (soft wet sounds, quiet moans, heels clicking on marble). For a spoken line, keep it to 5-10 words and prefix an emotion label in brackets: [breathlessly] "come closer" lip-syncs far better than a long sentence.
Two things the model does on its own when you say nothing: it adds a music track, and it invents dialogue. no background music in the tail kills the first. For the second, either write the line you want or write none at all, and keep the sound purely to what the scene itself makes.
Avoid
- Speed words - fast, quick, rapid, suddenly and "slow motion" all cause jitter and tearing. Say "steady" or "gently" instead.
- Re-describing the image - colors, species, outfit, background. Spend every word on motion, light and sound.
- Image tags - quality and artist tags such as masterpiece, 8k or by_artist do not help motion.
- Big multi-step actions - one clear move only, whatever the length.
- Words that read as age cues -
small,tiny,little,petite,diminutive, and pet names in a spoken line (baby,babe,kid,girl,boy). The content filter reads them as a minor and rejects the whole prompt, even on an obviously adult scene. Writeher frame,her body, or use the species instead. - Undressing or covering up - describing clothes coming off or going on is what makes an outfit flicker in and out across the clip.
How long it can be
Manual takes up to 1500 characters, which is far more than one action needs.
Tips
- Auto writes it for you. Reach for Manual when you want the exact line, or when you are continuing a clip.
- In Manual, your prompt is used as written. It goes to the model without rewriting, so write a complete line (action + light + sound), not a single word.
- Named copyrighted characters may be rejected by the video provider.
- For still-image prompts, see Positive Prompts.
Advanced: per-second control
For fine control you can narrate the clip second by second. These are beats inside one continuous take, not a shot list: never change angle between them, or you get the cut you were trying to avoid. Anchor the scene at At 00:00 (pose, light, framing), then describe only what changes each second, building intensity:
At 00:00 she kneels in soft warm side light, lips parted, hands on his thighs.
At 00:01 her head dips slightly, breath hitching.
At 00:02 the motion deepens, her ears flatten.
At 00:03 her eyes squeeze shut, muscles tensing.
At 00:04 she settles back into a slow steady rhythm.
static camera, single continuous shot, clothing unchanged, same art style, no extra characters or limbs, no background music, no censorship, stable face, no distortion, no jitter, clean loop.
Keep adjacent moments close (small changes) to avoid jitter, and stay well under the character limit.