Cued Speech-Integrated Audio-Visual Variational Autoencoder for Speech Enhancement
摘要
Speech enhancement (SE) is essential for improving the quality and intelligibility of speech signals, particularly in noisy environments. In this paper, we propose an innovative approach to Audio-Visual Speech Enhancement (AVSE) by modifying an audio-visual variational autoencoder (AV-VAE) framework to integrate both lip movements and hand gestures from Cued Speech (CS) as visual cues. This is the first work to incorporate hand gestures, in addition to lip movements, within the AVSE task. By introducing hand cues, our approach aims to address the inherent challenges of lip reading, such as the high ambiguity in interpreting lip movements, which can limit the effectiveness of traditional AVSE methods. Leveraging deep learning and computer vision techniques, our method offers a more comprehensive representation of spoken content. Through empirical evaluation, we demonstrate the effectiveness of our approach in enhancing the clarity and quality of speech signals, even in challenging acoustic conditions. The results indicate that the integration of hand cues significantly improves speech quality, providing a promising solution for AVSE in noisy environments.