Automatic Speech Recognition
摘要
Automatic Speech Recognition (ASR) is a complex technology that translates spoken language into written text. It involves audio signal analysis, acoustic modeling, pronunciation modeling, and language modeling. ASR is an expansive subject that cannot be thoroughly covered in a single chapter, obviously. This monograph aims to introduce the fundamental ASR concepts essential to understanding spoken language processing and the role ASR plays in this process. We begin by presenting the history of ASR systems. Next, we explore the basic notions relevant to ASR: speech production, acoustic modeling, pronunciation modeling, and language modeling, and we explain their roles in the overall architecture of the system. Decoding of speech is also described along with popular metrics used to evaluate the performance of ASR systems. We conclude the chapter with a brief discussion of possible future research directions in the space of ASR systems.