MaskTalker: Audio-Driven Talking Head Generation from Masked Face Using StyleGAN
摘要
Human mouths in the wild are often intentionally or unintentionally masked by obstacles, which hinders the lip reading for communication comprehension. To alleviate the issue, we present a novel approach called MaskTalker that generates talking head videos with lip motions, head poses, and eye blinks from the masked faces and audio. The task is innovatively formulated as seeking a trajectory navigating from the source to the target in the known latent space of StyleGAN conditioned on audio. However, two major challenges exist: the fuzzy direction forward due to coupling of current latent space, information scrambling and distortion caused by obstacles and high compression rate. To this end, we explore another LipSpace, which learns a set of directions exclusively for mouth shape transformation in a more disentangled way, clarifying the direction forward. Next, we propose a Refinement Network to compensate for the lost information. The module compares the generated coarse results with input images and employs a flow-based model to revise the distorted details, which serve as the auxiliary for fine-grained videos generation. Extensive experiments are performed to verify the effectiveness and superiority of our MaskTalker. We also present the applications of our system in real-world scenarios.