Advances in OpenASR21 Evaluation with Increased Temporal Resolution for Speech Self-supervised Learning Models
摘要
The OpenASR21 evaluation consisted of speech recognition for low resource languages in 3 evaluation conditions: constrained, contrained plus, and unconstrained. In this paper we investigate the constrained plus condition. In the constrained plus condition, we can use any self supervised learning (SSL) model to reduce the word error rate (WER). The idea was to get good speech recognition accuracy with only 10 h of acoustic training data for the 15 low resource languages in OpenASR21. In this paper, we show that we reduce WER for all the 15 languages when we increase the temporal resolution of feature parameters computed from the speech SSL models from 20 ms to 10 ms. The temporal resolution of the SSL models is in general 20 ms. This increase in temporal resolution is done without retraining the SSL models. The resulting feature parameters with increased temporal resolution lead to 3.9% average absolute reduction in WER (from 1.2% for Javanese to 7.8% for Amharic) for the development set of the 15 languages in the OpenASR21 evaluation. We also compare WER for 5 different pre-trained SSL models in the low resource OpenASR21 languages scenario.