The process of producing speech is complex and includes a number of biosignals in addition to acoustics. In order to overcome the limitations of conventional speech processing in particular and to gain a better understanding of the process of speech creation in general, these biosignals can be used. They originate from the articulators, the movement of the articulator muscles, the connections within the brain, and the brain itself. With an emphasis on speech production, recognition, and volitional control, we discuss artificial mouth techniques in this review that make use of a variety of sensors, including gyros, images, 3-axial magnetic sensors, electromyography as EMG, electroencephalography as EEG, electropalatography as EPG, electromagnetic articulography as EMA, permanent magnet articulography as PMA, and articulator electromyography. Before classifying them into taxonomy, we evaluate the flow of several voice recognition-related deep learning technologies, including visual speech recognition and silent speech interface. We conclude by talking about ways to address the communication issues that persons with speech impairments face as well as upcoming deep learning research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Overview of Automatic Speech Recognition Based on Deep Learning and Bio–Signal Sensors

  • N. Venkatesh,
  • K. Sai Krishna,
  • M. P. Geetha,
  • Megha R. Dave,
  • Dhiraj Kapila

摘要

The process of producing speech is complex and includes a number of biosignals in addition to acoustics. In order to overcome the limitations of conventional speech processing in particular and to gain a better understanding of the process of speech creation in general, these biosignals can be used. They originate from the articulators, the movement of the articulator muscles, the connections within the brain, and the brain itself. With an emphasis on speech production, recognition, and volitional control, we discuss artificial mouth techniques in this review that make use of a variety of sensors, including gyros, images, 3-axial magnetic sensors, electromyography as EMG, electroencephalography as EEG, electropalatography as EPG, electromagnetic articulography as EMA, permanent magnet articulography as PMA, and articulator electromyography. Before classifying them into taxonomy, we evaluate the flow of several voice recognition-related deep learning technologies, including visual speech recognition and silent speech interface. We conclude by talking about ways to address the communication issues that persons with speech impairments face as well as upcoming deep learning research.