Unsupervised 3D Face Reconstruction Method Based on ITV-Net
摘要
ITV-Net (ImageTransformer and VarEncoder) proposes a new approach to address the challenge of scarce ground truth data in 3D face reconstruction. By integrating Transformer and Variational Autoencoder (VAE) for encoding and decoding, and introducing noise perturbations in the latent space, the method enhances feature diversity and representation. Utilizing deep learning and the 3DMM model, it enables fast 3D face reconstruction. In perceptual loss, normal consistency and reflection losses are incorporated to constrain the geometric structure in 3D reconstruction and enhance lighting reflection accuracy in the projected 2D images. Experiments show that the method performs excellently in complex scenes, especially in cases with large pose variations.