In this paper, we address computer-aided speech diagnosis by designing a method for the automated detection of Polish sibilants in preschool children. Our database was recorded from 47 children aged four to seven using a 15-channel data acquisition device. We propose a modified YAMNet architecture to classify short speech segments from the main channel into four classes. The segments are represented by a dedicated acoustic image based on filter-bank energies and their derivatives. We use a set of time-series data augmentation procedures over the data from all microphones to improve training. With the segment classification results, we determine a frame-wise speech segmentation to extract sibilants. Our segment classification model yields overall accuracy of 87.9%, with the sibilant classification recall and precision at 92.3% and 87.7%, respectively. The sibilant segmentation accuracy reaches 96.2% with an F1 score of 73.5%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Segmentation of Polish Sibilants Using Modified YAMNet Architecture for Computer-Aided Speech Diagnosis in Children

  • Michał Kręcichwost,
  • Paweł Badura,
  • Artur Piet,
  • Abid Hasan,
  • Natalia Moćko,
  • Zuzanna Miodońska,
  • Agata Sage,
  • Marcin Grzegorzek

摘要

In this paper, we address computer-aided speech diagnosis by designing a method for the automated detection of Polish sibilants in preschool children. Our database was recorded from 47 children aged four to seven using a 15-channel data acquisition device. We propose a modified YAMNet architecture to classify short speech segments from the main channel into four classes. The segments are represented by a dedicated acoustic image based on filter-bank energies and their derivatives. We use a set of time-series data augmentation procedures over the data from all microphones to improve training. With the segment classification results, we determine a frame-wise speech segmentation to extract sibilants. Our segment classification model yields overall accuracy of 87.9%, with the sibilant classification recall and precision at 92.3% and 87.7%, respectively. The sibilant segmentation accuracy reaches 96.2% with an F1 score of 73.5%.