Speech Recognition-Based Human–Computer Interaction: A Survey
摘要
The survey comprehensively analyzes existing research on Automatic Speech Recognition (ASR) systems. The analysis begins with an in-depth examination of the current state of ASR technology, detailing the core components, methodologies, and performance benchmarks that define contemporary ASR systems. These include the various algorithms used for speech-to-text conversion, the integration of machine learning and deep learning techniques, and the role of large datasets in training robust ASR models. Following this overview, the discussion identifies emerging trends and technological advancements. Key areas of focus include the development of more accurate and efficient ASR models, integrating ASR with other technologies such as natural language processing (NLP) and artificial intelligence (AI), and the increasing importance of multilingual and domain-specific ASR systems. Additionally, the paper explores the challenges ahead, such as improving ASR performance in noisy environments, addressing privacy concerns, and reducing biases in ASR systems. Ultimately, this study aims to provide a clear and nuanced understanding of the advancements shaping the ASR landscape. It highlights significant contributions from researchers, identifies gaps in the current knowledge, and suggests potential directions for future research.