Generate video with AI from a prompt, a frame, or both.
The first decision is not which model — it is what you hand it. A sentence, a still you have already approved, two stills that fix the move, or references that carry a subject somewhere new. Same scene below, same model, one variable.
One written scene, run twice through one model. The only difference is that the second one was also handed the frame. The sentence asked for a mug; without the frame it came back without a handle.
Both clips are the same model at the same settings for the same five seconds. Nothing here is a comparison between vendors — it is a comparison between the four things you can start from.
Video generation is not one thing.
Two frames, and you have specified the move.
A prompt can describe a camera move; two frames decide it. This is the same mug and the same studio with a second still on the end — a close framing generated from the first one, so it is provably the same room — and the model has to get from one to the other in five seconds. What you are buying here is not a better clip, it is a clip that lands somewhere you already chose, which is the difference between a shot and a lucky take.
References move the subject, not the scene.
The same mug, on a café table it was never photographed on. This is the mode people miss: the reference is not a first frame, so the model is free to build a new scene around a subject it has to keep. That is the one that matters for a catalogue or a character — the thing stays itself while everything else changes. It is also the mode with the loosest grip, so check the details you care about rather than assuming they survived.
Every video model, behind one key.
Built for visual inference.
Serving video is a different problem from serving text — a single request can saturate a GPU, and none of the tricks that made language models cheap apply. Hedra's engine was built for exactly that workload.
Whichever model you choose — ours or anyone else’s — it runs on the same infrastructure, behind one key. Which is what makes the choice on this page cheap to make: run the same scene through two models, or through two starting points, and keep the one that worked. Switching is a string in a request, not a migration.
Generate with API or Agent
Build with the API
One key, every model and every starting point. Same request shape whichever you pick.
Generate one now.
No code — write the scene, or drop in the frame you have already approved.
FAQs
From an image, whenever you already have one you are happy with. A prompt has to describe a composition; a still simply is one, and the clips on this page are what that difference looks like. Start from a prompt when the shot does not exist yet and you are still exploring — then take the frame you like and animate that.
Image to video →