Classification of Cleft Lip and Palate Speech Using Fine-Tuned Transformer Pretrained Models
摘要
Cleft lip and palate speech (CLP) is a cranio-facial disorder which leads to spectro-temporal distortions in the speech of an individual. This makes accessibility of CLP speakers to speech enabled applications which require Human-computer interaction (HCI) such as voice assistants very challenging. Recently the availability of pretrained models have made the constraint of low resource language very convenient. Recent findings have proven that pretrained transformer models perform way ahead of traditional classifiers. In this paper, with an aim to achieve high end classification results, pretrained Transformer models fine-tuned on CLP data are used. The results obtained from the transformer models such as Wav2Vec2, SEW, SEW-D, UniSpeechSat, HuBERT, DistilHuBERT showed a comparative performance of the models and specially DistilHuBERT showed a significant improvement in the accuracy being close to 100%.