US20260195932
2026-07-09
Physics
G06T11/00
The patent application introduces a novel system for generating customized text-to-video outputs, focusing on high-quality videos that incorporate specified identities and motion patterns. It utilizes an appearance-agnostic motion learning approach to separate motion patterns from appearance features. A spatial-temporal collaborative composition scheme is employed to integrate learned multi-subject and motion LoRAs, enhancing the flexibility and generalizability of text-to-video generation.
The proposed system includes several components: a subject learner, a motion learner, and a spatial-temporal collaborative composer. The subject learner generates token low-rank adaptations (LoRAs) for each subject, while the motion learner creates a motion LoRA from reference motion videos. The spatial-temporal collaborative composer then combines these LoRAs to produce a coherent video output, integrating multiple subjects and motion patterns using spatial-temporal sampling.
The method involves receiving images of subjects, a motion video, and a text prompt. A reference motion is learned through an appearance-agnostic motion learning process, guided by the text prompt. A spatial-temporal algorithm is applied to compose an initial video, which is then used to generate the final output video featuring the specified subjects and motion patterns.
The system architecture comprises a receiver for input parameters, a LoRA generator, and processors executing code to create the output video. The LoRA generator produces subject and motion LoRAs using the input data, while the processors utilize a spatial-temporal collaborative composition algorithm within a diffusion model framework to synthesize the video, ensuring precise control over noise and motion integration.
The disclosed processes provide a unified framework for video content customization, enabling control over subject identities and motion patterns. By employing subject and motion LoRAs, and utilizing appearance-agnostic motion learning alongside spatial-temporal composition, the system enhances user control and coherence in video generation, supporting complex multi-subject interactions and dynamic motions.