Accurate labeling of tone sandhi and polyphones is essential when creating a high-quality speech corpus for building a Mandarin text-to-speech system. Proper tone labeling can ensure that the constructed text-to-speech system generates a natural prosody. This paper proposes an iterative method for tone labeling using a deep learning-based tone recognizer. The iterative method labels the tones of four subsets of linguistic units that would be incorrectly labeled by a linguistic processor or could not be exhaustively checked by linguistic expertise when the speech corpus is large. These subsets include Tone 3 Sandhi, yi/bu tone sandhi, polyphones, and characters or words that may have different phonetic realizations than lexical tones due to prosodic structure or speakers’ idiomatic use. In the first iteration, a tone recognizer (or a labeler) is constructed by considering the lexical tones that a linguistic processor labels as tone targets. In the later iteration, the syllables in the four subsets are re-labeled with the tones that the tone recognizer gives with the highest score under some constraints, and the tone recognizer is re-trained with the re-labeled tones until convergence is reached. The experimental results showed that the proposed method could robustly label tones for syllables of tone sandhi and polyphones on a multi-speaking rate Mandarin speech corpus. Furthermore, this study found that syllables misrecognized as different tones from lexical tones may reflect the actual tone realizations caused by coarticulation, location in a prosodic structure, and speaking rates. This study also provided a quantitative analysis of the relationship between labeled tones and prosodic structure to conform to the characteristics found in previous linguistic studies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tone Labeling for Mandarin Speech Corpus: A Deep Learning Iterative Approach

  • Chen-Yu Chiang,
  • Wu-Hao Li,
  • Te-Hsin Liu

摘要

Accurate labeling of tone sandhi and polyphones is essential when creating a high-quality speech corpus for building a Mandarin text-to-speech system. Proper tone labeling can ensure that the constructed text-to-speech system generates a natural prosody. This paper proposes an iterative method for tone labeling using a deep learning-based tone recognizer. The iterative method labels the tones of four subsets of linguistic units that would be incorrectly labeled by a linguistic processor or could not be exhaustively checked by linguistic expertise when the speech corpus is large. These subsets include Tone 3 Sandhi, yi/bu tone sandhi, polyphones, and characters or words that may have different phonetic realizations than lexical tones due to prosodic structure or speakers’ idiomatic use. In the first iteration, a tone recognizer (or a labeler) is constructed by considering the lexical tones that a linguistic processor labels as tone targets. In the later iteration, the syllables in the four subsets are re-labeled with the tones that the tone recognizer gives with the highest score under some constraints, and the tone recognizer is re-trained with the re-labeled tones until convergence is reached. The experimental results showed that the proposed method could robustly label tones for syllables of tone sandhi and polyphones on a multi-speaking rate Mandarin speech corpus. Furthermore, this study found that syllables misrecognized as different tones from lexical tones may reflect the actual tone realizations caused by coarticulation, location in a prosodic structure, and speaking rates. This study also provided a quantitative analysis of the relationship between labeled tones and prosodic structure to conform to the characteristics found in previous linguistic studies.