Abstract
In this paper, we propose a new single-photo 3D reconstruction model DiffuseVoxels focused on 3D inpainting of destroyed parts of a building. We use frustum-voxel model 3D reconstruction pipeline as a starting point for our research. Our main contribution is an iterative estimation of destroyed parts from a Gaussian noise inspired by diffusion models. Our input is twofold. Firstly, we mask the destroyed region in the input 2D image with a Gaussian noise. Secondly, we remove the noise through many iterations to improve the 3D reconstruction. The resulting model is represented as a semantic frustum voxel model, where each voxel represents the class of the reconstructed scene. Unlike classical voxel models, where each unit represents a cube, frustum voxel models divides the scene space into trapezium shaped units. Such approach allows us to keep the direct contour correspondence between the input 2D image, input 3D feature maps, and the output 3D frustum voxel model.