错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Arabic Lipreading Using YOLO and CNN Models

  • Ali Baaloul,
  • Nadjia Benblidia,
  • Abdelkader Ouared,
  • Fatma Zohra Reguieg

摘要

Lipreading is a vital aspect of human communication and requires effective computational methods. Lip movements, integral to this process, present challenges such as variability and context dependence. Recent developments in deep learning show potential for enhancing Arabic visual speech recognition (VSR) systems. This paper focuses on leveraging deep learning to assist Arabian individuals with hearing impairments, reduce their communication barriers, and enhance their quality of life. We employed our Arabic-created dataset, including YOLO version V7 as a frontend for mouth detection and CNN models (i) InceptionV3, and (ii) custom CNN model for speech classification. Our approach aims to address the complexities of lipreading. Our results show promise, with an impressive 90% speech recognition accuracy. These results underscore the capacity of deep learning to improve visual speech recognition, Facilitating the development of more effective and precise methods for detection and recognition.