错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating Lattice-Free Acoustic Modeling for Children Automatic Speech Recognition in Low-Resource Settings Under Mismatched Conditions

  • Virender Kadyan,
  • Puneet Bawa,
  • Richa Choudhary

摘要

The progress of Automatic Speech Recognition (ASR) for children has been slower in languages with limited resources due to various challenges such as lack of training data, differences in acoustics, and intrinsic peculiarities of native speakers. This study aims to investigate the use of Lattice-Free Maximum Mutual Information (LF-MMI) acoustic modeling for children in ASR low-resource settings with mismatched conditions. In this paper, the experimentations have been performed through the development of three different ASR systems under matched as well as mismatched conditions with an objective of assessing the performance of a children's test dataset employing the Deep Neural Network (DNN) methodology of acoustic models. In similar manner, the requirement for large training data has initially been met through an internal perturbation with evaluation on a Time Delay Neural Network (TDNN) approach with cross entropy (CE) alongside lattice free sequence discriminative training employing Maximum Mutual Information (LF-MMI) and boosted-Maximum Mutual Information (bMMI) objective functions. The study also addressed the improved methodology by mitigating acoustic mismatch existing between adult and children speakers through fine-tuning of adult speech based on two prosody modification parameters: pitch and duration scaling. Furthermore, the best output of the modified adult train dataset through speaker adaption based on in-domain training data augmentation technique has been utilized to expand the original training speech. The study also addressed acoustic mismatch and inter-speaker variations through speaker adaptation and Vocal Tract Length Normalization (VTLN)-based approaches. The results show that the implemented approaches significantly improved the ASR performance on the test dataset, achieving an overall Relative Improvement (RI) of 61.86% and demonstrating competitive model performance and efficiency. Overall, this study provides insights into addressing data scarcity and other challenges in ASR for children in low-resource settings with mismatched conditions.