The state-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems are mostly based on the data-driven methods. However, low-resource languages may lack data for training. Articulatory Features (AFs) describe the movements of the vocal organ which can be shared across languages. Thus, this paper investigates AFs-based semi-supervised techniques to share data between languages. First, the traditional acoustic features and the AFs are combined as front-end features to provide articulatory information for cross-lingual knowledge transfer. Then, the dropout-based lattice decoded are used as the pseudo-labels for the unsupervised data to address the problem of data deficiency. In addition, the Lattice-free Maximum Mutual Information (LF-MMI) objective is adopted to better adapt to small datasets. Experiments show that our system can obtain a relative improvement of 58.6% on Character Error Rate (CER) comparing to the baseline system. More specifically, the smaller the datasets are, the more obvious the advantages of our system can be.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semi-supervised Cross-Lingual Speech Recognition Exploiting Articulatory Features

  • Xinmei Su,
  • Xiang Xie,
  • Chenguang Hu,
  • Shu Wu,
  • Jing Wang

摘要

The state-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems are mostly based on the data-driven methods. However, low-resource languages may lack data for training. Articulatory Features (AFs) describe the movements of the vocal organ which can be shared across languages. Thus, this paper investigates AFs-based semi-supervised techniques to share data between languages. First, the traditional acoustic features and the AFs are combined as front-end features to provide articulatory information for cross-lingual knowledge transfer. Then, the dropout-based lattice decoded are used as the pseudo-labels for the unsupervised data to address the problem of data deficiency. In addition, the Lattice-free Maximum Mutual Information (LF-MMI) objective is adopted to better adapt to small datasets. Experiments show that our system can obtain a relative improvement of 58.6% on Character Error Rate (CER) comparing to the baseline system. More specifically, the smaller the datasets are, the more obvious the advantages of our system can be.