US20260212560
2026-07-23
Physics
G06T11/60
The disclosed methods and systems enhance novel view synthesis (NVS) using machine learning models, focusing on image-to-image diffusion processes. Techniques include inpainting-based synthesis using depth maps, video synthesis with texture map inpainting, and high-fidelity image diffusion via texture conditional models. Additionally, methods for generating training data to improve NVS model training are discussed, incorporating symmetry exploitation and alignment error simulations.
NVS aims to create new views of a scene from existing data, useful in generating stereo images and videos. This is particularly relevant in virtual reality (VR), where stereoscopic viewing enhances the 3D experience. However, converting single images or videos to stereo pairs can be challenging, often resulting in visual errors that disrupt the immersive experience. Addressing these challenges, the current methods aim to improve the quality and efficiency of NVS processes.
Existing NVS methods often produce images with visual artifacts, such as unrealistic blurring or mismatched textures. These issues arise from inadequate training data quality. Machine learning models require extensive and high-quality datasets to perform effectively, yet multi-view datasets are scarce and costly. Therefore, improving training data availability and quality is crucial for enhancing NVS model performance.
The methods propose generating training data from single-view datasets, which are more abundant than multi-view datasets. By utilizing depth estimation and warping techniques, single-view images can be transformed into training sets suitable for NVS model training. This approach broadens the range of available training data, enhancing model efficiency and output quality.
The training process involves creating pairs of masked and ground-truth images, enabling NVS models to learn inpainting techniques. Models are trained to reconstruct masked images, which are then compared to ground-truth images for accuracy. This systematic training approach ensures that models can generate plausible novel view images, addressing previous limitations in NVS techniques.