Breaking the Corpus Bottleneck for Multi-dialect Speech Recognition with Flexible Adapters
摘要
Multi-dialect speech recognition has gained significant attention in recent years. In the realm of multi-dialect speech recognition, predominant strategies involve multi-dialect joint learning and cross-dialect transfer learning. However, these methodologies frequently give rise to undesirable interference among distinct languages, with a more pronounced impact on low-resource languages. Adapter modules were recently introduced as an efficient alternative to fine-tuning in multilingual speech recognition, but it generates a large number of parameters and it is biased towards higher-resource languages, especially for some multilingual corpus with unbalanced data distribution. To tackle these issues, in this paper, we propose a novel approach to improve recognition quality with a few additional parameters. Specifically, the proposed model comprises a global adapter module, three language-specific adapter modules, and a gating unit, facilitating the acquisition of specific knowledge from different dialects. This design alleviates mutual interference and mitigates learning biases arising from imbalanced data. Additionally, through extensive ablation experiments, we substantiate that the proposed method achieves a significant reduction in Character Error Rate (CER) compared to the baseline method, all while maintaining much greater parameter efficiency.