A Comprehensive Analysis on Features and Performance Evaluation Metrics in Audio-Visual Voice Conversion
摘要
Audio-Visual Voice Conversion (AVVC) is an emerging research field within the realm of audio-visual speech synthesis, involving the transformation of both vocal characteristics and lip movements from a source speaker to a target speaker while preserving linguistic content. Unlike conventional Voice Conversion (VC), AVVC incorporates visual cues alongside speech features to facilitate cross-domain transformations. This technology is driven by advancements in deep learning (DL) algorithms which have supplanted traditional statistical methods in AVVC model enhancements. Despite these advancements, evaluating the quality of AVVC-generated audio and video samples remains a formidable challenge within the research community. This paper systematically analyzes the essential features employed in AVVC models, encompassing both spectral and prosodic attributes. Furthermore, the paper delves into the myriad performance evaluation metrics utilized for assessing the efficacy of these models, including subjective and objective measures. The critical examination of these metrics sheds light on their applicability in the context of audio-visual voice conversion, highlighting the challenges and considerations specific to this field. The extraction of features and analysis of performance evaluation metrics provides a holistic understanding of the challenges and opportunities in this emerging field, aiming to contribute to the advancement of AVVC technologies.