错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identity Preserved Expressive Talking Faces with Synchrony

  • Karumuri Meher Abhijeet,
  • Arshad Ali,
  • Prithwijit Guha

摘要

This work proposes a novel approach to talking face generation using driving audio. The driving audio and a single image of the target person are provided as input to the proposed model. The model generates a realistic video of the target person uttering the driving audio. Recent works in this domain have focused on either one of expressions or lip-sync or identity preservation. This model provides supervision over photo realism, expression fulfilment, identity preservation and audio-visual synchrony which are crucial factors in synthesizing a realistic video. The proposed system is end-to-end trainable and the learning is performed with six losses. This method can generate photo realistic, expressive and audio-synced talking faces while preserving the identity of the target person. This work proposes a discriminator network to impose audio-visual synchrony in the generated video. The proposed model is trained on RAVDESS dataset containing 24 professional actors (12 female and 12 male), uttering two statements in a neutral North American accent with disgust, sad, angry, happy, surprise, fearful and calm emotions. This work is benchmarked on the VID-TIMIT dataset against three baseline models.