Automatic Speech Recognition (ASR) is a complex technology that translates spoken language into written text. It involves audio signal analysis, acoustic modeling, pronunciation modeling, and language modeling. ASR is an expansive subject that cannot be thoroughly covered in a single chapter, obviously. This monograph aims to introduce the fundamental ASR concepts essential to understanding spoken language processing and the role ASR plays in this process. We begin by presenting the history of ASR systems. Next, we explore the basic notions relevant to ASR: speech production, acoustic modeling, pronunciation modeling, and language modeling, and we explain their roles in the overall architecture of the system. Decoding of speech is also described along with popular metrics used to evaluate the performance of ASR systems. We conclude the chapter with a brief discussion of possible future research directions in the space of ASR systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Speech Recognition

  • Mikołaj Morzy

摘要

Automatic Speech Recognition (ASR) is a complex technology that translates spoken language into written text. It involves audio signal analysis, acoustic modeling, pronunciation modeling, and language modeling. ASR is an expansive subject that cannot be thoroughly covered in a single chapter, obviously. This monograph aims to introduce the fundamental ASR concepts essential to understanding spoken language processing and the role ASR plays in this process. We begin by presenting the history of ASR systems. Next, we explore the basic notions relevant to ASR: speech production, acoustic modeling, pronunciation modeling, and language modeling, and we explain their roles in the overall architecture of the system. Decoding of speech is also described along with popular metrics used to evaluate the performance of ASR systems. We conclude the chapter with a brief discussion of possible future research directions in the space of ASR systems.