Hermosso

"Glittered skin, editorial light"

Generate with Hermosso

Hermosso

"Glittered skin, editorial light"

Generate with Hermosso

HomeGlossary › Text-to-video

Text-to-video

Text-to-video is generating a moving clip from a written description alone, with no input footage. The model invents the scene and the motion — and on newer engines, the sound.

Text-to-video extends image generation across time: the model must keep characters, objects and light consistent frame after frame, which is why the field matured years after still images. Current frontier engines produce clips of roughly 4 to 15 seconds; longer films are built by chaining shots rather than prompting one long take.

A concrete example: "a paper boat floating down a rain gutter at dawn, cinematic macro shot, reflections in the water" returns a five-second clip with camera drift, ripples and — on engines with native audio — the sound of rain.

How Hermosso uses it: Text to Video runs the same 18-engine picker as photo animation, from budget drafts to flagship quality, and the video models page compares them head to head.

Try it on your own photos

Upload a few selfies, and Hermosso trains a private AI model of you — then generates studio-quality photos in any style. Your first credits are free.

Create your AI photos →