Detection and Classification of Categories of Dysphonia Using Convolutional Neural Network
摘要
Recent advances in computer science have contributed to the emergence of algorithms for the automatic detection of vocal pathologies to complement medical evaluation. However, such algorithms only perform binary classifications between healthy and disordered voices or classify only the most recurrent dysphonias in the databases. This paper evaluates a new methodology for accomplishing this detection, using the etiology of the dysphonia as a classification criterion, which results in 3 categories: functional, organic, and organofunctional. According to this criterion, the voice signals present in the Saarbruecken Voice Database were classified. Then, the identification of these categories of dysphonia was performed by extracting spectrograms of vocal signals and classifying them with a Convolutional Neural Network (CNN). The results show that CNN successfully classified organic dysphonia, reaching an accuracy of 75.4%, while for functional dysphonia, it obtained 67.5%, and for the multi-label classification, 52.9%. Such results indicate that the proposed feature extraction method was efficient and can be improved to recognize all categories of dysphonia with high classification scores.