Scene headings survive the breakdown
A screenplay already encodes location, time of day and who is present. The shot breakdown reads that structure rather than paraphrasing your script into a prompt and losing it.
Script to video means two different things. Most tools take a paragraph and pick stock clips to sit under a voiceover. This one takes an actual screenplay, breaks it into numbered shots, casts it, and renders each shot — which is a slower and more expensive answer to a harder question.
영상A screenplay already encodes location, time of day and who is present. The shot breakdown reads that structure rather than paraphrasing your script into a prompt and losing it.
Every named character in the script gets a canonical image, and every shot they appear in is dispatched with that image attached. Continuity comes from the script's own dramatis personae.
You get a numbered table with camera notes and action lines. Merge shots, split them, cut them, rewrite the action — all at text prices, before a single frame is generated.
The dominant meaning of "script to video" in this market is script-to-stock: you paste a paragraph, the tool writes a voiceover, and it cuts matching clips from a royalty-free library. It is fast, it is genuinely cheap — Wondershare Filmora's version costs 50 credits, around thirty cents — and no frames are generated. Its own documentation describes selecting visuals from "a built-in library of royalty-free stock footage and images".
That is the right tool for an explainer or a listicle, and paying generative prices for it would be a mistake. It is the wrong tool the moment your script has characters in it, because the library does not contain your characters.
The other meaning — the one this page is about — is script-to-render. A screenplay goes in, comes out as a numbered shot list, gets cast, and each shot is generated. Nothing is drawn from a library. That costs dollars rather than cents per finished minute, and it is the only route to footage that matches a script nobody has shot yet.
Text-to-video answers "describe a shot and I will render it". One prompt, one clip. It is the right entry point when the idea is a single image in motion and you do not yet know what the story around it is.
Script-to-video answers "here is the whole thing, break it down and shoot it". The input is structured — scenes, characters, dialogue — and that structure is what makes the shot breakdown possible. You are not describing a shot; you are handing over a document and getting a production plan back.
The practical difference shows up in what you edit. On the text page you rewrite prompts. Here you rewrite the shot list: shot 7 is two shots, shot 12 should be a close-up not a wide, the scene at the harbour needs an establishing shot before the dialogue. Then you render.
Three script pages is roughly three minutes of screen time, which typically breaks down to twenty shots. At eight seconds a shot that is about 160 seconds of rendered video.
| Model | Per 8s shot | Twenty shots | Holds a cast? | Native audio |
|---|---|---|---|---|
| MiniMax H3 | 72 credits ($0.72) | 1,440 credits ($14.40) | 9 refs | Yes |
| PixVerse C1 | 79 credits ($0.79) | 1,568 credits ($15.68) | 7 refs | Yes |
| Grok Imagine Video | 85 credits ($0.85) | 1,696 credits ($16.96) | 7 refs | No |
| Kling O3 | 135 credits ($1.35) | 2,688 credits ($26.88) | 4 refs | Yes |
| Kling V3.0 | 152 credits ($1.52) | 3,040 credits ($30.40) | 4 refs | Yes |
| Vidu Q3 | 186 credits ($1.86) | 3,712 credits ($37.12) | 4 refs | Yes |
| Seedance 2.0 | 365 credits ($3.65) | 7,296 credits ($72.96) | 9 refs | Yes |
| Veo 3.1 | 480 credits ($4.80) | 9,600 credits ($96.00) | 3 refs | Yes |
| Seedance 2.5 | 568 credits ($5.68) | 11,360 credits ($113.60) | 30 refs | Yes |
Only the reference-capable models are listed, because a script with named characters needs them. Keyframes add 10 to 48 credits a shot depending on the image model; the script and breakdown stages are text and cost a rounding error.
Standard screenplay formatting is fine — scene headings, action, character cues and dialogue are all read. There is no page limit that matters before the render stage, because everything up to it is text.
Every named character gets a subject sheet. Upload a photo instead if you have someone specific in mind; an upload outranks a generated sheet everywhere downstream.
This is where directing happens. Add coverage, cut what does not earn its seconds, set shot lengths — length is what you will be billed for, so this is also where the budget is set.
One still per shot to check framing and likeness, then video. Shots concatenate in order; the stream is copied through untouched, so nothing is re-encoded or stamped.
Reelfo does not stamp a watermark on anything it renders. The compose step copies the video stream through untouched — there is no overlay filter in the pipeline to apply one.
That is worth stating plainly because "script to video without watermark" is a heavily searched phrase, and in most of this market it describes a paid unlock. Here there is nothing to unlock: no overlay stage exists, so free output and paid output are the same file.
Signing up puts 100 credits in your account, once, with no card. That is enough for the script, character and shot-list stages plus about one keyframe — it is a look at the pipeline, not a free finished reel.
For a twenty-shot script that means the grant covers the breakdown, the cast and a keyframe — the whole plan — and the render is the part you pay for. Which is the right way round: you find out whether the shot list is any good before spending anything on motion.