US20260162278
2026-06-11
Physics
G06T7/215
The patent application introduces a system designed to optimize the processing of multimodal foundation models, which handle video/image and text data. By segmenting video frames into patches and generating tokens for these patches, the system uses motion information to classify them as either motion or no-motion. Tokens from no-motion patches are then pruned based on system conditions like power or temperature, reducing computational demands while maintaining accuracy in tasks such as object detection.
Multimodal foundation models are advanced AI systems capable of processing diverse types of input data, including images, videos, text, and audio. These models, such as vision language models (VLMs) and vision language action models (VLAMs), integrate language models with non-text data encoders to produce outputs ranging from video analytics to actionable commands. The integration allows these models to understand and process multimodal data effectively.
The system employs image token pruning to address challenges posed by large image token sizes, which can increase computational costs. By leveraging motion information, such as motion vectors, the system identifies and prunes redundant no-motion image tokens at various layers of the model. This technique reduces the data processed, thereby lowering computational requirements without affecting the accuracy of the model's predictions.
Image token pruning offers significant benefits in terms of reduced latency, memory, and power usage, enabling real-time AI applications in sectors like security, retail, and robotics. The reduction in computational load enhances performance per watt and cost efficiency, allowing existing hardware to handle more processing tasks without needing additional components. This is particularly beneficial in edge deployments with limited resources and environmental constraints.
The patent describes an edge server environment where motion-based pruning circuitry operates. The server processes video streams from cameras, utilizing decoder and pre-process circuitry to prepare data for multimodal foundation models. These models generate video analytics, supported by object tracking and post-process circuitry, to deliver actionable insights and alerts. This setup illustrates the practical application of the pruning technique in improving system efficiency and performance.