Masked Appearance Restoration for High Resolution Talking Face Generation
摘要
Guidance-free talking face generation is a complex and challenging task for high-resolution videos. Most guidance-free methods generate the target video by deforming the appearance of reference frames to match the source video. However, the potential pose discrepancy between reference and source frames often leads to appearance distortions in these methods. To address this issue, we propose a Masked Appearance Restoration (MAR) approach for generating high-resolution videos with reasonable facial structure and detailed texture. Specifically, the proposed MAR comprises two main modules: a Masked Structure Reasoning (MSR) module and a Motion Detail Aligning (MDA) module. The MSR module focuses on reasoning multi-scale structure features by combining a masked autoencoder with multi-scale upsampling layers, which ensures that the generated videos exhibit precise structural representations. Subsequently, to preserve detailed textures, the MDA module leverages audio-driven cross-attention to dynamically align the texture and structure features. Finally, the structure and texture features are fused and fed into a decoder, generating the target video frames. Extensive experiments indicate that our method generates accurately lip-synced talking faces in high-resolution video. Moreover, our model demonstrates robustness to extreme pose variations.