<p>We present a system for hearing-impaired people, designed to enhance the accessibility of video podcasts by identifying and displaying the current speaker. Current methods often focus on either audio–visual integration or speech transcription to identify the speaker. As a result, the provided solutions are partial and without comprehensive integration. Departing from the state-of-the-art, this paper introduces the ECHO model, combining speech embeddings using the Wav2Vec2 model and clustering methods like agglomerative clustering, for precise speaker identification with synchronized subtitle generation. The proposed approach includes lip movement analysis to identify the current speaker using the MTCNN face detection model and the LipNet lip detection model. Additionally, audio is isolated from video sources, while the pre-trained Whisper model is used for speech transcription. Speaker identification is achieved through speech encoding and clustering, facilitating the discernment of distinct speakers within the audio–visual context. For pristine empirical evidence, a benchmark dataset of talk shows for speaker identification is introduced. The proposed model demonstrates significant improvements, achieving at least 6% gains over the state-of-the-art on the proposed dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ECHO: enhanced communication for hearing impaired in online video podcasts

  • Gouthami Godhala,
  • Vijayasree Asam,
  • Samriddha Sanyal

摘要

We present a system for hearing-impaired people, designed to enhance the accessibility of video podcasts by identifying and displaying the current speaker. Current methods often focus on either audio–visual integration or speech transcription to identify the speaker. As a result, the provided solutions are partial and without comprehensive integration. Departing from the state-of-the-art, this paper introduces the ECHO model, combining speech embeddings using the Wav2Vec2 model and clustering methods like agglomerative clustering, for precise speaker identification with synchronized subtitle generation. The proposed approach includes lip movement analysis to identify the current speaker using the MTCNN face detection model and the LipNet lip detection model. Additionally, audio is isolated from video sources, while the pre-trained Whisper model is used for speech transcription. Speaker identification is achieved through speech encoding and clustering, facilitating the discernment of distinct speakers within the audio–visual context. For pristine empirical evidence, a benchmark dataset of talk shows for speaker identification is introduced. The proposed model demonstrates significant improvements, achieving at least 6% gains over the state-of-the-art on the proposed dataset.