Datasets are fundamental for training speech translation models and play an important role in advancing research in this field. Currently, high-quality Tibetan-Chinese speech translation datasets are relatively scarce, hindering research progress in related languages. To address the scarcity of Tibetan-Chinese speech translation data, this paper utilizes speech synthesis technology to create a Tibetan-Chinese speech translation dataset for the three major Tibetan dialects: Amdo, Kham, and Ü-Tsang. First, Tibetan text data is cleaned based on an open-source Tibetan-Chinese bilingual parallel sentence pair dataset. Then, the VITS model is used to generate speech data for the three major Tibetan dialects, which is processed and reviewed by experts. After correcting and modifying the dataset, we obtain a multi-dialect speech translation dataset from Tibetan dialect speech to Chinese text. This dataset includes Amdo Tibetan speech, Ü-Tsang Tibetan speech, Kham Tibetan speech, and corresponding Tibetan and Chinese texts, containing a total of 8,979 samples, totaling 2.3 GB and 10.74 h. Among these, Amdo Tibetan audio data comprises 4.28 h, Kham Tibetan audio data comprises 3.56 h, and Ü-Tsang Tibetan audio data comprises 2.90 h, with an audio sampling rate of 16kHz. This dataset can be used for research on translating Tibetan dialect speech to Chinese text, as well as for research on Tibetan dialect speech recognition and Tibetan speech anti-spoofing, providing data support to advance research in related fields.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-dialect Tibetan-Chinese Speech Translation Dataset Based on Speech Synthesis

  • Qing Lin,
  • Liping Zhu

摘要

Datasets are fundamental for training speech translation models and play an important role in advancing research in this field. Currently, high-quality Tibetan-Chinese speech translation datasets are relatively scarce, hindering research progress in related languages. To address the scarcity of Tibetan-Chinese speech translation data, this paper utilizes speech synthesis technology to create a Tibetan-Chinese speech translation dataset for the three major Tibetan dialects: Amdo, Kham, and Ü-Tsang. First, Tibetan text data is cleaned based on an open-source Tibetan-Chinese bilingual parallel sentence pair dataset. Then, the VITS model is used to generate speech data for the three major Tibetan dialects, which is processed and reviewed by experts. After correcting and modifying the dataset, we obtain a multi-dialect speech translation dataset from Tibetan dialect speech to Chinese text. This dataset includes Amdo Tibetan speech, Ü-Tsang Tibetan speech, Kham Tibetan speech, and corresponding Tibetan and Chinese texts, containing a total of 8,979 samples, totaling 2.3 GB and 10.74 h. Among these, Amdo Tibetan audio data comprises 4.28 h, Kham Tibetan audio data comprises 3.56 h, and Ü-Tsang Tibetan audio data comprises 2.90 h, with an audio sampling rate of 16kHz. This dataset can be used for research on translating Tibetan dialect speech to Chinese text, as well as for research on Tibetan dialect speech recognition and Tibetan speech anti-spoofing, providing data support to advance research in related fields.