US20260260396
2026-09-03
Physics
G06T11/00
The invention introduces a method for creating video clips from text descriptions using a diffusion model. The process begins by receiving a text description and converting it into a vector representation. This vector is then used to generate a sequence of key points for the video, which are synthesized by a diffusion motion model. These key points are mapped to images representing each frame of the video, culminating in the generation of a complete video clip sequence.
This method is situated within the realm of machine learning, particularly focusing on models that transform text prompts into video clips. The approach leverages synthetic information to define the motion dynamics of objects within the video frames, ensuring the generated video aligns with the concepts in the text description.
Current generative models can convert text descriptions into video clips, but often lack quality and precision in aligning with the text's concepts. Some models enhance video quality by using reference video images, but finding suitable references is challenging due to limited availability. This invention aims to overcome these limitations by generating high-quality video clips without relying on reference videos, thus broadening the scope of possible concepts that can be represented.
The process involves several key steps: receiving a text description, obtaining its vector representation, and using this to generate a sequence of key points through a diffusion motion model. These key points are then mapped to images corresponding to video frames, resulting in a coherent video clip. The method is implemented in an electronic device equipped with processors and memory that execute the necessary instructions to perform these operations.
Illustrations and diagrams accompany the detailed description of the process, showcasing the sequence of operations and system components. Figures depict the diffusion process, training data generation, and the mapping of key points to images. The electronic device is configured to execute the method, ensuring the seamless transformation of text descriptions into dynamic video clips, as demonstrated in the provided examples.