Six stages, not one prompt box
Story, characters, shots, keyframes, video, final cut. Each stage is a revision you can read and edit before the next one spends anything, so a bad shot list gets fixed for free instead of getting rendered.
An AI short film generator from text has one hard job that a clip generator does not: the person in shot one has to still be that person in shot six. Reelfo writes the script, casts the characters, breaks the story into shots, renders a keyframe for each, and only then generates video — carrying a locked character reference through every stage.
影片Story, characters, shots, keyframes, video, final cut. Each stage is a revision you can read and edit before the next one spends anything, so a bad shot list gets fixed for free instead of getting rendered.
Characters are generated once as a subject sheet, then that image is passed as a reference into every shot that features them. Shot six reads the same reference shot one did — it is not a fresh roll of the dice.
Only models with a real reference-image field appear on the reel surface. The ones that read a first frame and then improvise are still on /video for single clips, but they never get handed a multi-shot story.
Ask any text-to-video model for a five-shot story and you will get five good-looking clips of five different people. The model has no memory between generations. Each call is a fresh sample from the same prompt, and "a woman in her thirties, dark hair, red coat" describes a very large set of faces.
The fix is not a better prompt. It is a pipeline that renders the character once, as an image, and then hands that exact image to the video model as a reference on every subsequent shot. That only works on models that expose a reference-image field — and most do not.
Reelfo publishes 9 video models on the reel surface out of 12 in the catalogue. The difference is not quality. Hailuo 2.3, Wan 2.6, HappyHorse 1.0 are perfectly good models — they simply take a starting frame and nothing else, so by shot two they are improvising a face. HappyHorse 1.0 is the clearest case: it reads the first frame, and the drift is visible in the second shot.
Reference slots is the number of distinct character images a single shot can be given — a two-hander needs at least two. Everything in this table is pulled from the live model registry, so it moves when the registry moves.
| Model | Reference slots | Clip length | Max resolution | Price / 5s |
|---|---|---|---|---|
| MiniMax H3 | 9 | 5–15s | 768P | 45 cr ($0.45) |
| PixVerse C1 | 7 | 1–15s | 1080P | 49 cr ($0.49) |
| Grok Imagine Video | 7 | 4–10s | 720P | 53 cr ($0.53) |
| Kling O3 | 4 | 3–15s | 84 cr ($0.84) | |
| Kling V3.0 | 4 | 3–15s | 95 cr ($0.95) | |
| Vidu Q3 | 4 | 1–16s | 1080P | 116 cr ($1.16) |
| Seedance 2.0 | 9 | 4–15s | 1080P | 228 cr ($2.28) |
| Veo 3.1 | 3 | 8–8s | 4K | 300 cr ($3.00) |
| Seedance 2.5 | 30 | 4–30s | 720P | 355 cr ($3.55) |
Credit prices are what Reelfo charges; 1 credit = $0.01. A shot longer than five seconds scales linearly from this rate.
Every stage writes a revision. You can rewrite any of them and only the stages downstream of your edit are re-run — changing a line of dialogue does not re-render the whole film.
You give a brief — one line is enough. The model returns a premise, a beat structure and a script. This stage is text only, so iterating on it costs almost nothing.
Each character in the script gets a subject sheet: one canonical image, generated once. You can also upload a photo instead, and an uploaded reference outranks a generated one everywhere downstream.
The script is broken into numbered shots, each with a camera note, an action line and the list of characters present. This is the storyboard, in table form, and it is fully editable.
One still per shot, rendered with the character references attached, so you can see the framing and the cast before committing to video. Models that bind references directly skip this stage rather than charging you for an image they will not read.
Each shot is dispatched to the chosen video model with its keyframe and its character references. This is the only stage with a confirmation gate, because it is the only one that costs real money.
The shots are concatenated in order with audio mixed in. The video stream is copied through untouched — no re-encode, no overlay, no watermark.
Credits are a flat $0.01 each, and video is billed per second of output at the model's own rate. A five-shot short at eight seconds a shot works out to 392 credits ($3.92) on PixVerse C1, 672 credits ($6.72) on Kling O3, or 1,824 credits ($18.24) on Seedance 2.0.
The text stages are effectively free next to that, and the keyframes are cheap: GPT Image 2 at 10 credits an image, Seedream 5.0 Pro at 43 credits an image, Nano Banana Pro at 48 credits an image. Which is why the pipeline puts every decision you might change in front of the video stage rather than behind it.
Signing up puts 100 credits in your account, once, with no card. That is enough for the script, character and shot-list stages plus about one keyframe — it is a look at the pipeline, not a free finished reel.
You have to sign in. There is no anonymous mode — generation is gated at the API, the tRPC layer and the render service, so "no sign up" is not something this tool can offer you. If that is a dealbreaker, it is better to know now than after writing a brief.
Reelfo does not stamp a watermark on anything it renders. The compose step copies the video stream through untouched — there is no overlay filter in the pipeline to apply one.
There is no desktop app to download. Everything runs in the browser and the finished MP4 downloads from there, on any plan, at full quality.
Individual clips top out at the per-model limits in the table above. A "short film" here means a sequence of shots cut together, not one continuous long take — no current model will hold a single unbroken minute.