A StyleCLIP-Based Facial Emotion Manipulation Method for Discrepant Emotion Transitions
摘要
Leveraging StyleCLIP’s expressivity and its disentangled latent codes, current methodologies enable facial emotion manipulation through textual inputs. Despite these advancements, significant challenges remain in manipulating target emotions that deviate markedly from the originals, especially in transitioning from happy to sad emotions without introducing artifacts or errors. This paper introduces a novel approach for discrepant emotion transitions. Our network architecture integrates a StyleGAN2 generator with an Emotion Manipulation Mapper, a Dual Auxiliary Classifier, and a CLIP Text Encoder. By utilizing the inverse cumulative distribution function, we convert source emotion labels into conditional data, thus enhancing the model’s ability to accurately map and modify the emotional distribution across faces. We evaluated our method against established techniques using the Radboud Faces Database and CelebA-HQ dataset, and introduced a new quantitative measure including seven metric for assessing manipulation efficacy.