Text-to-Video Generation
Models that generate video clips from a text prompt or a reference image — used for ads, social content, and increasingly for avatar-presenter (talking head) video.
This category splits into two distinct use cases that are easy to conflate: general text-to-video (generating novel scenes/motion from a prompt, still limited in clip length and consistency) and avatar/presenter video (a realistic AI presenter speaking a script, used heavily for training and marketing content, which is a more mature and reliable capability today than general scene generation). Evaluation criteria differ accordingly — clip length, motion consistency, and prompt adherence matter for the former; lip-sync quality, voice cloning, and language coverage matter for the latter. As with generative image models, commercial usage rights and likeness/consent policies (especially for avatar video) are a real diligence item, not boilerplate.
Last verified: