Communication occurs commonly between two or more individuals and often it involves the use of speech. However, under certain circumstances, the use of speech may be restricted causing a hindrance in communication. These circumstances may include noisy environments, people with hearing disability, and corrupt audio in forensic analysis. The alternative to speech in this situation can be to visually analyze the content from the patterns of lip movements to deduce the message. This phenomenon is known as lip-reading. Lip-reading can be intricate using the naked eye for an individual. Artificial Intelligence (AI) has gained recent attention in many fields, with speech recognition being one of them. The use of Machine Learning (ML) and Deep Learning (DL) algorithms, falling under the umbrella of Artificial Intelligence, can help build Automatic Lip-Reading (ALR) systems that can ease the process of lip-reading. This work aims to build an ALR system while obtaining a great performance compared to some of the state-of-the-art methods on the Grid corpus dataset. The proposed model is segregated into two parts: frontend and backend. The front end employs a 3D Convolutional Neural Network (CNN) and the back end uses a Bi-directional Long Short-Term Memory (Bi-LSTM) architecture. The model attained a Word Error Rate (WER) of 0.093, outperforming several baseline models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Efficient Approach to Lip-Reading with 3D CNN and Bi-LSTM Fusion Model

  • Rohit Chandra Joshi,
  • Aayush Juyal,
  • Vishal Jain,
  • Saumya Chaturvedi

摘要

Communication occurs commonly between two or more individuals and often it involves the use of speech. However, under certain circumstances, the use of speech may be restricted causing a hindrance in communication. These circumstances may include noisy environments, people with hearing disability, and corrupt audio in forensic analysis. The alternative to speech in this situation can be to visually analyze the content from the patterns of lip movements to deduce the message. This phenomenon is known as lip-reading. Lip-reading can be intricate using the naked eye for an individual. Artificial Intelligence (AI) has gained recent attention in many fields, with speech recognition being one of them. The use of Machine Learning (ML) and Deep Learning (DL) algorithms, falling under the umbrella of Artificial Intelligence, can help build Automatic Lip-Reading (ALR) systems that can ease the process of lip-reading. This work aims to build an ALR system while obtaining a great performance compared to some of the state-of-the-art methods on the Grid corpus dataset. The proposed model is segregated into two parts: frontend and backend. The front end employs a 3D Convolutional Neural Network (CNN) and the back end uses a Bi-directional Long Short-Term Memory (Bi-LSTM) architecture. The model attained a Word Error Rate (WER) of 0.093, outperforming several baseline models.