Start with a photo of your face, add your script and a short clip of your voice — and get a lip-synced talking video that sounds like you. No camera, studio or editing needed. Talking Photo turns a still photo into a talking video; Lipsync Video takes a short clip of you talking and makes it say new lines, with the voice pulled from the video itself.
Upload a photo of your face, a short clip of your voice (6+ seconds) and your script. We clone your voice and animate the photo so it speaks your lines with synchronized lip movement. If you already have a video of yourself talking, Lipsync mode reuses the voice from that video instead.
For Talking Photo, yes — a 6+ second clip of clear speech is how the video ends up sounding like you. For Lipsync Video, no separate recording is needed: we pull the voice straight from the video you upload.
Scripts are capped at 75 words, which works out to about 30 seconds of speech — sized for social posts, updates and short explainers.
Videos are billed per second of output: 100 credits/sec for Standard 480p or 200 credits/sec for Premium 720p Talking Photo, and 100–150 credits/sec for Lipsync. New accounts start with 750 free credits — no card required — and credits never expire.
What you upload is used for one thing only: creating and running your personal AI model and generating your content. We never use your photos or voice to train models for other users, and we never sell or share them for advertising. Your trained model is private to your account.
Open the video studio, read the guide talking AI videos, or see plans and pricing.