Phone-Based Speech Recognition for Phonetic E-Learning System
摘要
The APTgt system is a web-based phonetics educational training tool that aims to improve teachers’ linguistics pedagogical experience and provide phonetic transcription training for students. Based on this system, we proposed an Automatic Speech Recognition (ASR) system to convert speech to a stream of phones, automatically generating the phonetic transcription of both standard speech and disordered speech. With the help of the phone recognizer, the instructors can generate a large number of phonetic transcription exams without manual transcription. This phone-level ASR system applied Mel-frequency cepstral coefficients (MFCCs) as features and bidirectional Long-Short Term Memory (LSTM) as an encoder. The Speech Exemplar and Evaluation Database (SEED) data set, including disordered and non-disordered speech, was used for training and further testing. The proposed recognizer will make our phonetic E-learning system more intelligent and better serve students with their performance on phonetic transcription.