<p>Automatic Speech Recognition (ASR) has been the regnant research area in the domain of Natural Language Processing for the last few decades. Past years’ advancement provides progress in this area of research. The accent of spoken is the prominent factor affecting the performance of ASR. Accent is the pronunciation of a word in a distinct way which may vary from specimen to specimen depending upon their social class, formality, age, vocal property, gender, geography, and influence of native or other language. Recognizing an accent requires various complex characteristics and features of voice such as voice quality, prosody, and phoneme pronunciation which are difficult to extract and analyze. To solve these difficulties, the researchers focus their attention on spectral representation such as Mel Spectrogram and Mel-Frequency Cepstral Coefficient (MFCC). In this study, we will analyze which features achieve maximum accuracy for the English accent classification task. Here in this work, we perform our experiments on Log Mel filter bank and MFCC features with four different window function. We extract features from raw audio with various windowing function and train a hybrid model of Convolutional Neural Network and Bidirectional Long Short-term Memory network (CNN-BiLSTM) on a custom dataset having nine different accents of English language namely American, Australian, British, Indian (Oriya, Bangla, Telegu, Malayalam), and Welsh. The accuracy of each feature is evaluated and compared. The log Mel filter bank outputs the highest accuracy of 99.75% and 99.91% of training and validation with Han window function while MFCC with Bartlett window function achieves 99.98% of training and 99.33% of validation accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing English accent identification in automatic speech recognition using spectral features and hybrid CNN-BiLSTM model

  • Ghayas Ahmed,
  • Aadil Ahmad Lawaye,
  • Vishal Jain,
  • Jyotir Moy Chatterjee,
  • Shubham Mahajan

摘要

Automatic Speech Recognition (ASR) has been the regnant research area in the domain of Natural Language Processing for the last few decades. Past years’ advancement provides progress in this area of research. The accent of spoken is the prominent factor affecting the performance of ASR. Accent is the pronunciation of a word in a distinct way which may vary from specimen to specimen depending upon their social class, formality, age, vocal property, gender, geography, and influence of native or other language. Recognizing an accent requires various complex characteristics and features of voice such as voice quality, prosody, and phoneme pronunciation which are difficult to extract and analyze. To solve these difficulties, the researchers focus their attention on spectral representation such as Mel Spectrogram and Mel-Frequency Cepstral Coefficient (MFCC). In this study, we will analyze which features achieve maximum accuracy for the English accent classification task. Here in this work, we perform our experiments on Log Mel filter bank and MFCC features with four different window function. We extract features from raw audio with various windowing function and train a hybrid model of Convolutional Neural Network and Bidirectional Long Short-term Memory network (CNN-BiLSTM) on a custom dataset having nine different accents of English language namely American, Australian, British, Indian (Oriya, Bangla, Telegu, Malayalam), and Welsh. The accuracy of each feature is evaluated and compared. The log Mel filter bank outputs the highest accuracy of 99.75% and 99.91% of training and validation with Han window function while MFCC with Bartlett window function achieves 99.98% of training and 99.33% of validation accuracy.