Evaluating the Degradation from Short Utterances and the Efficacy of In-domain Data Augmentation Techniques in Children’s ASV Systems
摘要
Automatic speaker verification (ASV) is a critical component of voice-based authentication systems and an essential task in ensuring the security and privacy of various voice-based applications. However, when it comes to children, the dearth of domain-specific data, developmental changes in vocal characteristics and hence a variability in children’s speech characteristics poses unique challenges for reliable and accurate verification systems. The difficulty increases even more with the presence of short-speech segments. This study explores the effectiveness of various in-domain data augmentation techniques, namely speed modification, pitch modification and vocal- tract length perturbation (VTLP); individually as well as in combination with each other. Speed and pitch perturbations enrich the diversity of acoustic characteristics by directly modifying the speech signal waveform, whereas the VTLP method facilitates data augmentation through adjustments at the feature level. The proposed approach of combining speed perturbation, pitch variation and vocal-tract length adjustments provide a more accurate estimation of the model parameters by not only expanding the training data-set but also accounting for a broader spectrum of children’s acoustic characteristics leading to a substantial reduction in equal error rate and false acceptance rates for children speaker verification tasks. This is demonstrated by a