Performance Analysis of MFCC and wav2vec on Stuttering Data
摘要
The development of Automatic Speech Recognition systems brings a need for automatic detection and identification of various disorders in speech. Stuttering is one of these types of disorders. The problem of stuttering detection and identification is challenging since existing datasets are either small in size or highly imbalanced. In this work, we conduct several experiments on a subset of the SEP28k dataset comparing two different feature representations for audio data.