Hierarchical Multi-task Learning with Articulatory Attributes for Cross-Lingual Phoneme Recognition
摘要
This chapter proposes a novel hierarchical multi-task architecture for multilingual phoneme recognition in unseen, low-resource languages. We optimize articulatory attribute and phoneme classifiers jointly and provide attribute probability distributions as an additional input to the phoneme classifier. To improve this hierarchical connection, we propose removing blank logits of the Connectionist Temporal Classification to compute more consistent attribute probability distributions. As the acoustic model, we use a hybrid convolution-transformer architecture. The models were trained and evaluated on a subset of 24 languages from the Mozilla Common Voice corpus. Regular or unmodified hierarchical multi-task learning did not improve phoneme error rates (PERs) compared to a phoneme-only baseline in our experiments. In contrast, the evaluation shows that the proposed blank removal method outperforms the baseline by 3.16 percentage points (pp.) PER. The evaluation of zero-shot cross-lingual transfer on a dataset with 95 languages shows negative transfer effects between attribute and phoneme classifiers in regular and unmodified hierarchical multi-task learning. However, using the novel blank removal approach, the hierarchical multi-task model achieved an absolute PER improvement of 2.69 pp. over the baseline.