Deep Auscultation with Demographic Data: Detecting Respiratory Anomalies Using Convolutional Neural Networks and Autoencoders
摘要
Digital auscultation, integrating digital signal processing and machine learning algorithms, has garnered significant attention in the field of automated disease diagnosis due to its simplicity, speed, and non-invasiveness. In our research, we introduce a multi-modal hybrid model that merges a wavegram logmel CNN, initially trained on the extensive audio dataset, with demographic data processed through an autoencoder. Additionally, we analyze the impact of two encoding methods on demographic information and integrating different numbers of snapshot models in the snapshot ensemble method on the model scores. The experimental results demonstrate that both the wavegram logmel CNN and multi-modal hybrid model with two encoding methods exhibit improvements of 2.1 \(\%\) , 3.3 \(\%\) , and 2.8 \(\%\) , respectively, compared to a single model when utilizing four snapshot models. Through a tenfold cross-validation on the ICBHI dataset, the multi-modal hybrid model with one-hot encoding achieves a remarkable model score of 82.5 \(\%\) in the four-classification task, surpassing previous research outcomes and achieving a 1.2 \(\%\) enhancement over the wavegram logmel CNN model, which solely employs respiratory cycles as input.