US20260203533
2026-07-16
Physics
G06F40/58
The patent application describes a system designed to enhance video accessibility for deaf and hard-of-hearing individuals by integrating sign language into video streams. This system involves both server-side and client-side components. The server-side processor utilizes natural language processing (NLP) and artificial intelligence (AI) models to convert subtitle text into a standardized transcription representation. Meanwhile, the client-side processor features a video player, gesture translator, and avatar rendering engine to animate a customizable avatar that performs sign language gestures in real-time with the video.
With the rise of video streaming as a primary mode of media consumption, accessibility remains a critical issue for users who are deaf or hard-of-hearing. Traditional subtitles often fail to capture the full expressiveness of sign language, while existing solutions like Picture-in-Picture (PiP) video tracks of interpreters have limitations in terms of compatibility and customization. This system addresses these challenges by providing an integrated, customizable sign language solution that enhances the viewing experience without significant bandwidth or compatibility issues.
The system comprises a server with a memory for storing video and subtitle data, executing instructions via NLP and AI models. These models generate an intermediate representation from the subtitle text, which is then converted into a standardized transcription representation, such as the Hamburg Notation System. This transcription is packaged into a sign language subtitle package with timing information and delivered to the client. The client, equipped with a video player and an avatar rendering engine, converts the transcription into a signing gesture language description, animating an avatar to perform sign language gestures synchronized with the video.
The method involves receiving a video, audio, and subtitle text at the server, where the NLP model generates an intermediate representation. This is then translated into a standardized sign language transcription using the AI model. The resulting sign language subtitle package is sent to the client, which uses it to animate an avatar in real-time. The avatar's speed is synchronized with the video using timing information, and customization options are available based on user or provider preferences. The sign language subtitle package can be generated on-demand or as part of a video workflow.
The system offers several advantages, including reduced bandwidth and storage costs compared to traditional video tracks of interpreters. The text-based sign language subtitle track is compatible with various video players and streaming technologies, enabling easy updates and error corrections. By leveraging AI and NLP, the system provides a scalable solution to improve video accessibility, offering a more expressive and nuanced viewing experience for users who rely on sign language.