Controllable Talking Head Synthesis by Equivariant Data Augmentation for Spatial Coordinates
摘要
Traditional talking head synthesis algorithms decompose the lips and head pose information from the facial landmark points based on the spatial point registration. In this paper, we show that the hypothesis of the traditional point registration methods is too strong to result in unnatural talking head. Instead of registration, we propose a latent lip-head pose coding method. The proposed method performs self-supervised learning and equivariant data augmentation to the facial landmark points. The experimental results show that the proposed latent lip-head pose coding method outperforms the traditional registration-based methods and can generate natural looking talking head with accurate mouth shapes.