Cued Speech (CS) is an innovative communication system that enhances lip reading by incorporating hand gestures. This method significantly improves comprehension for individuals with hearing impairments. The development of automatic CS recognition and generation facilitates more effective communication between deaf individuals and the hearing community. The CS dataset is essential for establishing an AI-based automatic recognition and generation model for CS. Previous CS datasets were mainly in English and French, and the data volume was small with a single cuer (i.e., people who perform CS), which hinders research progress in this field. Therefore, we have constructed, for the first time, a Mandarin Chinese CS Dataset (MCCSD) containing 4000 CS videos from four native Chinese CS cuers. Importantly, we propose a novel GAN-based CS Video Gesture Generation baseline for the first time. To further validate the effectiveness of this dataset, we build a benchmark for both automatic CS video recognition and generation. Experimental results demonstrate that MCCS serves as a valuable benchmark for CS recognition and generation, presenting new challenges and insights for future research. The complete dataset, benchmark, and source codes, will be made publicly available.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCCS: The First Open Multi-Cuer Mandarin Chinese Cued Speech Dataset and Benchmark

  • Li Liu,
  • Lufei Gao,
  • Wentao Lei,
  • Yuzhi He,
  • Yuxing He,
  • Che Feng,
  • Yue Chen,
  • Zheyu Li

摘要

Cued Speech (CS) is an innovative communication system that enhances lip reading by incorporating hand gestures. This method significantly improves comprehension for individuals with hearing impairments. The development of automatic CS recognition and generation facilitates more effective communication between deaf individuals and the hearing community. The CS dataset is essential for establishing an AI-based automatic recognition and generation model for CS. Previous CS datasets were mainly in English and French, and the data volume was small with a single cuer (i.e., people who perform CS), which hinders research progress in this field. Therefore, we have constructed, for the first time, a Mandarin Chinese CS Dataset (MCCSD) containing 4000 CS videos from four native Chinese CS cuers. Importantly, we propose a novel GAN-based CS Video Gesture Generation baseline for the first time. To further validate the effectiveness of this dataset, we build a benchmark for both automatic CS video recognition and generation. Experimental results demonstrate that MCCS serves as a valuable benchmark for CS recognition and generation, presenting new challenges and insights for future research. The complete dataset, benchmark, and source codes, will be made publicly available.