Text-to-video
Text-to-video is generating a moving clip from a written description alone, with no input footage. The model invents the scene and the motion — and on newer engines, the sound.
Text-to-video extends image generation across time: the model must keep characters, objects and light consistent frame after frame, which is why the field matured years after still images. Current frontier engines produce clips of roughly 4 to 15 seconds; longer films are built by chaining shots rather than prompting one long take.
A concrete example: "a paper boat floating down a rain gutter at dawn, cinematic macro shot, reflections in the water" returns a five-second clip with camera drift, ripples and — on engines with native audio — the sound of rain.
How Hermosso uses it: Text to Video runs the same 18-engine picker as photo animation, from budget drafts to flagship quality, and the video models page compares them head to head.
Try it on your own photos
Upload a few selfies, and Hermosso trains a private AI model of you — then generates studio-quality photos in any style. Your first credits are free.
Create your AI photos →