Invention Title:

REAL-TIME AUDIO-BASED FULL-BODY GESTURE GENERATION FOR 3D AVATAR

Publication number:

US20260203983

Publication date:
Section:

Physics

Class:

G06T13/205

Inventors:

Applicant:

Smart overview of the Invention

The patent application describes a method for generating real-time 3D avatar animations based on audio input. The process involves buffering audio streams and extracting speech features using an audio encoder. These features are then fed into a gesture generation model, which predicts the 3D positions and orientations of body joints, enabling the creation of realistic avatar animations. The animations are synchronized with the audio playback, enhancing user engagement through lifelike gestures.

Technical Background

The development responds to market demand for on-device, conversational 3D AI avatars that integrate with large language models (LLMs). Current animation technologies primarily focus on facial expressions, lacking real-time full-body gesture generation. This innovation addresses these limitations, providing a more immersive user experience by adding full-body movements to avatars.

Methodology

The invention buffers audio streams, extracts speech features, and inputs them into a gesture generation model. This model predicts the 3D position of the body's center and joint orientations, which are used to generate animation keyframes. These keyframes are stored in memory, and the audio plays in sync with the avatar's gestures. Users can select avatar modes and motion blending options to tailor the animation style.

Embodiments

The patent outlines multiple embodiments, including a method, an electronic device, and a non-transitory machine-readable medium. Each embodiment facilitates real-time avatar animation through sequential audio buffering, speech feature extraction, and gesture prediction. The animations can be adjusted for head-only or full-body movements, with options for blending raw and predetermined animations.

Customization and Deployment

Users can customize avatar gestures through style inputs and motion blending modes, allowing control over hand motions and animation intensity. The technology can be deployed on various devices, including mobile, extended reality (XR), and robots. This flexibility ensures broad application across different platforms, enhancing interactive experiences with realistic 3D avatars.