US20260229215
2026-08-06
Physics
G10L13/02
A device captures an image of application data displayed to a user on an electronic display. It then generates a prompt for a generative artificial intelligence model to summarize the captured image. This summary is sent back to the device, which facilitates reading it aloud to the user, enhancing accessibility for visually impaired users navigating dynamic applications.
The focus is on improving accessibility for visually impaired users interacting with dynamic applications. Existing tools like "Job Access With Speech" (JAWS) struggle with dynamic content and interactive visuals, creating barriers for users. This innovation addresses these limitations by leveraging AI to provide real-time summaries of visual data.
Visually impaired users face challenges with computer screens, often relying on screen readers that perform poorly with dynamic content. Accessibility tools depend on websites implementing specific features, which is not always feasible due to resource constraints. This invention seeks to overcome these challenges by using AI to interpret and communicate complex visual information.
The device captures images of application data and uses AI to generate summaries. This process involves a computing system with client devices, servers, and databases interconnected via networks. The system supports various communication protocols and can be part of cloud-based services, enhancing the flexibility and reach of the solution.
The computing system comprises client devices, servers, and databases, communicating over networks using both wired and wireless connections. The device includes network interfaces, processors, and memory to execute software processes, including AI and accessibility functions. This architecture supports diverse devices, from desktops to IoT devices, ensuring broad applicability.