HMM-GMM Acoustic Modeling for Arabic Speech Recognition System
摘要
Acoustic model is a fundamental component in the construction of a speech recognition system, serving a critical function in converting raw audio signals into phonetic units that can be accurately interpreted and processed by the system. This paper presents the development of an acoustic model tailored for Arabic speech recognition system, which integrates a Hidden Markov Model (HMM) with Gaussian Mixture Models (GMM). The HMM effectively captures the temporal dynamics of speech, while the GMM computes the emission probabilities of observed acoustic features based on the hidden states. To enhance the feature representation of the model, Mel-frequency cepstral coefficients (MFCCs) are utilized alongside their first and second derivatives, capturing both the static and dynamic characteristics of the speech signal. For training and evaluation the model, the Arabic Sound Database (ASD) has been created, consisting of 34 distinct phoneme classes and includes a total of 272000 audio files. An extensive experimental evaluation was conducted to identify the optimal architectural parameters, revealing that a configuration consisting of 7 states and 16 GMM mixtures achieved the highest performance, with an accuracy of 91.47%. These findings demonstrate the efficacy of the HMM-GMM model in accurately identifying the complex acoustic patterns that are distinctive of Arabic speech.