错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Study of Pre-trained Audio and Speech Models for Heart Sound Detection

  • Yuxin Duan,
  • Chenyu Yang,
  • Zihan Zhao,
  • Yiyang Jiang,
  • Yanfeng Wang,
  • Yu Wang

摘要

Cardiovascular disease screening is critically anchored in heart sound auscultation. As deep learning methodologies advance, the impetus toward automating heart sound detection grows, aiming to curtail reliance on specialized clinicians. However, the compilation and annotation of expansive high-fidelity datasets present challenges, attributing to both necessary expertise and environmental complexities. In this landscape, transfer learning, harnessing extensive pre-trained models, emerges as a potential solution. In our investigation, we rigorously assessed established audio and speech models—PANNs, SSAST, BEATs, HuBERT, and WavLM—using the PhysioNet/CinC 2016 dataset. Preliminary results showcased the pre-tuning BEATs model’s superior performance, achieving an accuracy of approximately 90%. However, following optimization procedures, the PANN-V1 model surpassed its counterparts, registering an accuracy of 94.02%. Our study further delved into the models’ robustness against various noise paradigms. Pink noise was observed to be more disruptive than white noise, with the PANN-V2 model demonstrating notable resilience across both noise spectra. Contrarily, impulse noise exhibited a minimal perturbative effect. In a more pragmatic setting, we evaluated the models using the CirCor DigiScope Dataset, emphasizing specific demographics such as pediatric and antenatal populations. It was discerned that these particular demographics, coupled with ambient clinical noise, can indeed modulate model performance. Within this context, the BEATs model retained commendable proficiency, achieving a 65.33% accuracy. This study provides insights into model selection and fine-tuning, fostering more informed decision in the selection of pre-trained models for heart sound processing and analysis.