Enhancing child-machine interaction for Indian children speaking English as non-native language using hybrid CNN and customized dictionary
摘要
The spoken language expression has found a new application with the advent of sophisticated conversational agents present in different set up of homes, offices as well as educational centers. The progress in state-of-the-art techniques of speech recognition have provided seamless solutions for adults to converse effectively with voice-bots or AI enabled virtual speech assistants. But the same systems struggle to interpret children’s speech effectively, particularly if the speech is non-native and belongs to a low resource language group. The major cause for this variation is the pronunciation and accent of the non-native language which the system is not equipped to handle in addition to the physical and temporal growth factors which the children go through in their growing years. Furthermore, the lack of publicly available datasets for children’s speech, speaking English as secondary language, adds to the complexity of communication with smart assistants. This issue of lack of children’s speech corpora is addressed in this work by explicitly creating a dataset of children in the age group of 5–15 years, having Hindi or Marathi as their native language and speaking English as non-native language. Our investigation and experimentation show an accuracy of above 95% which is potentially better than experiments done on similar low-resource Indian languages. The contribution of this work is an end-to-end pipeline that covers creating an annotated speech sample dataset of Indian children, extraction of MFCC and i-vector features from the corpora, training and evaluating the hybrid Convolutional Neural Network (CNN) on non-native speech, employing a language model and dictionary that can capture the variation in pronunciation and checking accuracy and performance of the system. The results have shown an improvement in the recognition of Indian children’s speech compared to standard speech to text translators.