Invention Title:

METHOD AND SYSTEM FOR GENERATING EMOTIONAL TALKING HEAD VIDEO USING DISENTANGLED POSE AND EXPRESSION FLOW GUIDANCE

Publication number:

US20260220864

Publication date:
Section:

Physics

Class:

G06T13/40

Inventors:

Assignee:

Applicant:

Smart overview of the Invention

The disclosed method and system focus on generating emotional talking head videos by employing disentangled pose and expression flow guidance. This process involves creating realistic animations of human faces that maintain identity, synchronize lip movements, and express emotions accurately. The system takes multiple inputs, including an identity image, speech audio, and emotion data, to produce these videos. The method utilizes distinct pose and expression generation networks to handle head movements and facial expressions separately, ensuring the final output is both realistic and expressive.

Technical Challenges

Generating emotional talking head videos is challenging due to the need for realistic head movements and emotional expressions that preserve the subject's identity. Traditional methods often require additional input videos to drive these elements, making them impractical for real-world applications. Existing methods also struggle with accurately capturing emotions due to the limited variability in available emotional datasets. Moreover, they often fail to generalize well to arbitrary faces and backgrounds.

System Components

The system comprises several networks, including a pose generation network based on conditional VAE-LSTM, and an expression generation network utilizing graph convolution. The pose generation network creates diverse head movements, while the expression generation network focuses on facial expressions. These networks work together to produce a coherent video by processing expression-invariant pose landmarks and pose-invariant emotion landmarks. The image generation network then synthesizes the final video using disentangled optical flow computation for both pose and expression guidance.

Methodology

The method begins by receiving inputs such as an identity image, speech audio, and emotion data. The pose generation network uses these inputs to generate head movements, while the expression generation network creates corresponding facial expressions. These are combined in the image generation network, which employs a disentangled flow approach to ensure the final video accurately reflects both the intended emotions and head movements. This approach allows for the creation of realistic and expressive talking head videos without relying on additional driving videos.

Applications and Benefits

This technology is crucial for enhancing user experiences in various commercial applications, such as digital assistants and virtual instructors. By overcoming the limitations of existing datasets and methods, the system can generalize better to different faces and backgrounds, providing more dynamic and lifelike interactions. The ability to generate talking heads that convey diverse emotions accurately is essential for creating engaging and effective human-computer interactions.