Text to Video
Text to video is the job of generating footage from a written description alone, with no photo or clip to start from. You write a prompt naming the subject, the action, the camera and the setting; the model invents every frame. It is the least predictable thing in this studio and the one where the prompt does all the work.
Five models here can do it, priced from 20 credits for Grok Video 1.5 to 320 for Seedance 2.5. The editor on the left shows the exact cost for the model and settings you pick. This route is not connected yet — the controls and prices are real, the Create button is not.
What a text-to-video prompt has to contain
A text-to-video prompt has to supply five things the model has no other way of knowing: subject, action, camera, setting and light. Leave one out and the model picks for you, which is where most disappointing results come from — not from a bad model, but from a prompt that only specified two of the five.
Write it as one flowing sentence rather than a list of tags. These models were trained on video captions, which are prose, not on comma-separated keyword strings borrowed from image generators.
| Element | Weak | Specific |
|---|---|---|
| Subject | a woman | a woman in a grey wool coat, mid-thirties |
| Action | walking | walking towards the camera, hands in her pockets |
| Camera | (missing) | a slow tracking shot moving backwards ahead of her |
| Setting | a street | a narrow street of wet cobblestones after rain |
| Light | (missing) | late afternoon, low sun behind her, long shadows |
Why the first clip is usually wrong, and what to do about it
Text to video has no anchor. Image to video at least starts from a picture you chose, so the look is settled before generation begins; here the model decides the face, the wardrobe, the lens and the colour grade all at once, and the odds of it landing on yours the first time are low.
The workable method is to treat the first generation as a draft of the prompt rather than a draft of the video. Generate on the cheapest model that can hold the scene, read what came back against what you asked for, and fix the element the model chose for you. Only move up to an expensive model once the prompt reliably produces the right shot cheaply.
Repeat the same prompt on a better model and you will get the same composition rendered better. Change the prompt and the model at the same time and you will not know which change did what.
Picking a model for text to video
The price range here is sixteen-fold, and it does not map neatly onto quality for every shot. A static scene with one subject looks close to identical across most of these models; a complex camera move through a crowded space is where the expensive ones pull ahead.
| Model | Per clip | Reach for it when |
|---|---|---|
| Grok Video 1.5 | 20 credits | You are still working out the prompt |
| Wan 3.0 | 45 credits | Everyday shots, one subject, simple camera |
| MiniMax H3 | 116 credits | The shot needs real camera work, or you want generated audio |
| Seedance 2 | 160 credits | A cinematic look matters and 2.5 is more than you need |
| Seedance 2.5 | 320 credits | The clip is going somewhere people will look closely |
Text to video or image to video?
Use image to video whenever you can produce or find the first frame, which is more often than people assume. Generating a still is cheap, fast and repeatable — you can iterate on a picture until the composition, face and wardrobe are right, then hand the finished frame to the video model and spend its budget entirely on motion.
Text to video earns its place when the shot cannot be photographed or generated as a still first: an aerial move over terrain, a shot that has to start mid-action, or a scene you are exploring before you know what it looks like.
When not to use this
- Getting a specific person, product or logo on screen. The model invents the subject every time. If it has to be a particular thing, start from a picture of that thing in image to video.
- Matching a shot you already have. There is no way to tell a text-to-video model to continue an existing clip from this page.
- Consistent characters across several clips. Each generation re-imagines the subject; the same prompt run twice gives you two different people.
- Readable text in frame. Signs, titles and labels come out as plausible-looking nonsense. Add them in an editor.
- Long shots. Every model here tops out around ten seconds, and the last seconds are the least stable.
- Anything depicting a real, identifiable person. Naming someone in a prompt to generate footage of them is not what this is for.
Questions people ask
What is text to video?
Text to video is a model that generates video frames from a written description, with no image or clip as input. You supply a prompt describing the subject, action, camera movement, setting and lighting, and it invents every frame of the result.
How do I write a good text-to-video prompt?
Name five things in one prose sentence: who or what the subject is, what it does, how the camera moves, where the scene is, and what the light is doing. Anything you leave out, the model chooses for you. Avoid comma-separated tag lists and words like "cinematic" or "8K" — those are image-generator habits and carry no information here.
How much does text to video cost?
Between 20 and 320 credits per clip depending on the model: Grok Video 1.5 at 20, Wan 3.0 at 45, MiniMax H3 at 116, Seedance 2 at 160 and Seedance 2.5 at 320. The editor shows the exact figure on the Create button before you spend anything.
Which text-to-video model is best?
For a single subject in a simple scene, the cheap ones are close enough that the price difference is hard to justify. The expensive models separate themselves on complex camera movement, crowded scenes and anything where several things move independently. Work out your prompt on Grok Video 1.5, then re-run the finished prompt on a better model.
Can text to video make a video of a specific person?
No, and you should not try. The model has no way to reproduce a particular individual from a name, and generating footage that depicts a real identifiable person is not something we support. If you need a specific subject, photograph it and use image to video.
Is text to video free?
No. Every generation costs credits because every generation costs us money upstream. The cheapest route to a usable clip is to work out the prompt on the 20-credit model rather than to look for a free tier.
Does the video have sound?
Only on MiniMax H3, which generates its own audio track. The other models return silent MP4s.
Related tools
Checked 20 September 2026.