US20260195645
2026-07-09
Physics
G06N20/00
The patent application describes a method for training a multi-modal base model using a processor-implemented approach. This involves generating a token data set from a target document's information, creating data chunks, and summarizing these chunks to train the model. The process integrates diverse data forms, such as text and images, enhancing the model's ability to understand and process multi-modal data, similar to human cognition.
The field of this invention relates to neural network-based models, specifically focusing on training with multi-modal data sets. These data sets combine different types of data, including text, images, voice, and video. The goal is to mimic human-like cognitive abilities, allowing models to interpret and integrate various data forms simultaneously, enhancing their understanding and functionality.
A key aspect of the method involves generating a token data set from a target document, which is then used to create initial data chunks. These chunks are summarized to produce summary token data, which is used to generate subsequent chunks. The training of the multi-modal base model occurs using these chunks, ensuring that the model can process and integrate different data types effectively.
An electronic device is described, equipped with processing circuitry and memory to store and execute instructions. These instructions enable the device to generate token data sets and data chunks, create summary token data, and train the multi-modal model. The method ensures that data chunks adhere to specified length constraints, optimizing the processing and training efficiency.
The application also highlights the ability to generate global summary token data, which aids in the creation of data chunks and enhances the model's training. The method supports the integration of intersecting image and text data from target documents, utilizing image encoders to process visual information. This comprehensive approach aims to improve the model's multi-modal processing capabilities, aligning with the detailed description's flexibility and adaptability.