US20260237131
2026-08-13
Physics
G06T13/00
The system introduces an AI-based framework for generating customized storytelling videos using a multi-agent approach. This framework includes five specialized agents: a story designer, a storyboard generator, a video creator, an agent manager, and an observer. These agents collaboratively interpret a user's textual prompt and reference video to produce a multi-shot video tailored to the subject of the reference video. The story designer, agent manager, and observer utilize Large Language Models (LLMs) to enhance narrative quality, while the storyboard generator and video creator ensure visual consistency through advanced AI techniques.
The invention addresses limitations in current automated video generation systems, which often struggle with maintaining high-quality, coherent storytelling. Existing methods like Sparse Control and Sparse Video Diffusion have attempted to improve video generation by focusing on narrative alignment and animation, yet they fall short in preserving subject consistency across frames. This system seeks to fill the gap by integrating advanced AI models to deliver videos with consistent character details and narrative flow, crucial for applications across education, entertainment, and marketing.
The system's architecture is designed to ensure efficient and high-quality video production. The story designer AI agent creates detailed storylines based on user inputs, which the storyboard generator translates into precise storyboards. The video creator then synthesizes these storyboards into videos, ensuring intra-shot consistency through a Latent Diffusion Model (LDM)-based Image-to-Video (I2V) generation model. The agent manager coordinates the agents' tasks, while the observer provides feedback to refine the final output, ensuring the generated videos meet the desired quality and narrative intent.
Key innovations include the use of a three-step pipeline by the storyboard generator to maintain character detail consistency across video shots, involving generation, removal, and redrawing steps. The system also incorporates LoRA-BE (Low-Rank Adaptation with Block-wise Embeddings) to enhance temporal consistency within shots. This multi-agent framework provides users with greater control and flexibility, allowing for fine-grained customization and adaptability to diverse storytelling inputs.
The system's versatility makes it suitable for a wide range of applications beyond customized storytelling video generation, addressing the growing demand for high-quality, adaptable video content in various domains. Its ability to maintain both inter-shot and intra-shot consistency, while adapting to different narrative requirements, positions it as a transformative solution in the field of AI-driven video production, offering significant improvements over existing methodologies.