Image to video
Animate an image you own into a five-second clip with synchronised audio.
Image to video turns a still image into a five-second clip with sound. It needs an active subscription, and it only works on images you own.
Starting a render
The Generate Video button appears on a gallery image when two things are true: the image is yours, and it doesn't already have a video. You'll find it in two places.
- Under an AI-generated image in a chat, on the row of actions with Edit image and Delete Image.
- On a character's gallery, on any image you own.
You own an image if you created the character it belongs to, or if you generated the image yourself. There is no entry point from moments, and an avatar is only animatable through the gallery.
Each generation costs 120,000 credits. The charge lands when the finished video is attached to the gallery item, so a render that fails costs nothing — but Generate Again is a second full-price generation, not a retry of the first.
The AI Video Director
The dialog is a shot planner, not a list of presets.
Design directions, or skip straight to writing
Design 3 directions sends the image to an AI pass that studies it and proposes three distinct five-second shots — each with a title, a subject motion, a camera move, an intensity and a sound direction. Pick one and it fills in the controls; Redesign asks for three more.
This is an extra AI call, billed on its own token use and waived while your subscriber allowance covers it. Writing Subject motion yourself skips it entirely.
Direct the motion
Set what moves, how strongly, and what the camera does.
Direct the sound
Choose a sound mode, and — if you want the character to talk — type the exact words.
Generate
The render streams its progress through directing, rendering and uploading.
Motion controls
| Control | Options |
|---|---|
| Subject motion | Free text, up to 1,200 characters |
| Motion intensity | Subtle · Natural · Expressive |
| Camera motion | AI chooses · Locked camera · Slow push in · Slow pull out · Pan left · Pan right · Tilt up · Tilt down · Gentle handheld |
| Output framing | Original · Portrait 3:4 · Landscape 16:9 · Vertical 9:16 · Square 1:1 |
| Optional ending frame | Another image from the same character's gallery |
Picking an ending frame changes how the camera behaves: the start-to-end transition takes over the camera path, so the clip lands on the second image.
Sound
The model generates audio inside the video, synchronised with what happens on screen — ambience, foley, music and character vocal sounds. It is not a soundtrack added afterwards.
| Sound design | What you get |
|---|---|
| AI chooses | The model picks the audio direction from the image and the motion |
| Ambience & foley | Environment and the sounds of what moves |
| Background music | Music under the shot |
| Character voice | The character makes vocal sound — and speaks your dialogue if you write any |
| Silent | No audio |
Four text fields sit under the mode, and each appears only when it can do something:
| Field | Limit | Shown when |
|---|---|---|
| Sound direction | 500 characters | The mode is anything but Silent |
| Exact dialogue (optional) | 500 characters | The mode is AI chooses or Character voice |
| Voice tone | 160 characters | You have typed dialogue |
| Spoken language | 80 characters | You have typed dialogue |
Leave Exact dialogue empty and the model will not invent lines — that is what the placeholder means by avoiding invented dialogue.
What comes out
| Property | Value |
|---|---|
| Duration | Exactly 5 seconds |
| Frame rate | 24 fps |
| Format | MP4 with an audio track |
| Original framing | Keeps the source aspect at roughly a 1280 × 720 pixel area, rounded to a 32-pixel grid |
| Portrait 3:4 | 832 × 1120 |
| Landscape 16:9 | 1248 × 704 |
| Vertical 9:16 | 704 × 1248 |
| Square 1:1 | 960 × 960 |
Every clip is five seconds. There is no length setting.
Adult content
The video step does not add a content filter on top of the image you started from. The dialog says as much in its own pitch: "Unrestricted generation for eligible adult content".
While it renders
One render at a time, per account. Starting a second while the first is going returns "Another video is still rendering. Wait for it to finish before starting a new one."
Closing the dialog does not cancel anything: "Generation continues in the background. Reopen this window to check on it." There is no notification when it finishes — reopen the image to see it. The server gives up on a render after 10 minutes.
The finished video
The video is attached to that one gallery image. The image action row gains Show video / Show image to flip between them, and anyone who can see the gallery image can play the video. Only you get the generate and delete actions.
Deleting a video is permanent. It removes the clip, deletes the stored file and restores the still image. Nothing is refunded, and animating the image again is a new 120,000-credit render.
Related
- Generating images — where the source image comes from
- Subscriptions — the plan requirement for video
- Credits — where the 120,000 lands
- Snapshots — publishing gallery content to a character's page