Human–Machine Interface (HMI) through Automatic Speech Recognition (ASR) is a system that enables the interaction between humans and machines through software applications. Speech Recognition (SR) is an interdisciplinary branch of computer science and computational linguistics that creates methods and tools to allow computers to recognize and translate spoken language into text with the primary advantage of being able to search. Most of today’s market is captured by human–machine interface devices, which use the concept of automatic speech recognition. Several researchers have developed ASR systems for different Indian regional and foreign languages, as well as for various dialects, using different techniques and tools. This study presents the different methods used by different researchers in the field of automatic speech recognition. It gives comparisons between different feature extraction methods and classification methods used in human–machine interfaces. It also explores the different challenges in this field. This study gives knowledge of speech recognition for isolated and connected words and continuous and spontaneous speech. This study will help other researchers develop technology-oriented solutions for human–machine interfaces. This study gives detailed information on how different researchers used different feature extraction methods and classification methods for different languages and it is observed that Mel Frequency Cepstrum Coefficients (MFCC) of feature extraction and Convolutional Neural Network (CNN) of classification methods give higher accuracy. The improvements and accuracy rates in ASR for both isolated and continuous speech in a variety of languages are highlighted in this paper.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Study and Analysis of Automatic Speech Recognition Systems for Indian and Foreign Languages

  • Sagar P. Chitte,
  • Madhukar N. Shelar,
  • Sunil S. Nimbhore,
  • Aashish R. Lahase

摘要

Human–Machine Interface (HMI) through Automatic Speech Recognition (ASR) is a system that enables the interaction between humans and machines through software applications. Speech Recognition (SR) is an interdisciplinary branch of computer science and computational linguistics that creates methods and tools to allow computers to recognize and translate spoken language into text with the primary advantage of being able to search. Most of today’s market is captured by human–machine interface devices, which use the concept of automatic speech recognition. Several researchers have developed ASR systems for different Indian regional and foreign languages, as well as for various dialects, using different techniques and tools. This study presents the different methods used by different researchers in the field of automatic speech recognition. It gives comparisons between different feature extraction methods and classification methods used in human–machine interfaces. It also explores the different challenges in this field. This study gives knowledge of speech recognition for isolated and connected words and continuous and spontaneous speech. This study will help other researchers develop technology-oriented solutions for human–machine interfaces. This study gives detailed information on how different researchers used different feature extraction methods and classification methods for different languages and it is observed that Mel Frequency Cepstrum Coefficients (MFCC) of feature extraction and Convolutional Neural Network (CNN) of classification methods give higher accuracy. The improvements and accuracy rates in ASR for both isolated and continuous speech in a variety of languages are highlighted in this paper.