Lung diseases have had a significant impact in recent years, according to the World Health Organization (WHO). From 2019 to 2023, deaths ranging from 1.3 to 3.23 million worldwide have been recorded, placing lung diseases in fourth place as cause of deaths worldwide. In this work, we propose using convolutional neural networks (CNN) to classify lung sound into four classes: normal, crackle, wheeze and both (crackle-wheeze). We evaluated the CNN using four different audio transforms applied over the input data: STFT, CQT, Mel, and Wavelet. The suggested CNN model is based on MobileNetV1 and uses regularization and dropout techniques to prevent overfitting. We also compared the results of the CNN to those obtained using AutoML as a tool to improve the neural network’s hyperparameters. According to the findings, the WT performs better in classifying lung sounds than other transforms with a score of 0.503, exhibiting higher accuracy, sensitivity and specificity for a CNN; while the STFT lags slightly behind the WT regarding classification performance with a score of 0.489. On the other hand, AutoML shows that the Mel transform with a score of 0.404 outperforms the STFT with a score of 0.395. These results suggest that the choice of the transform to apply over the input data, significantly influences the accuracy of the classification of lung sounds.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Audio Transformation Techniques for Pulmonary Sound Classification Using CNN and AutoML

  • Christian Alejandro Saldaña Calderón,
  • Said Polanco Martagón,
  • Luis Antonio González Castro,
  • Yahir Hernández Mier,
  • Marco Aurelio Nuño Maganda

摘要

Lung diseases have had a significant impact in recent years, according to the World Health Organization (WHO). From 2019 to 2023, deaths ranging from 1.3 to 3.23 million worldwide have been recorded, placing lung diseases in fourth place as cause of deaths worldwide. In this work, we propose using convolutional neural networks (CNN) to classify lung sound into four classes: normal, crackle, wheeze and both (crackle-wheeze). We evaluated the CNN using four different audio transforms applied over the input data: STFT, CQT, Mel, and Wavelet. The suggested CNN model is based on MobileNetV1 and uses regularization and dropout techniques to prevent overfitting. We also compared the results of the CNN to those obtained using AutoML as a tool to improve the neural network’s hyperparameters. According to the findings, the WT performs better in classifying lung sounds than other transforms with a score of 0.503, exhibiting higher accuracy, sensitivity and specificity for a CNN; while the STFT lags slightly behind the WT regarding classification performance with a score of 0.489. On the other hand, AutoML shows that the Mel transform with a score of 0.404 outperforms the STFT with a score of 0.395. These results suggest that the choice of the transform to apply over the input data, significantly influences the accuracy of the classification of lung sounds.