<p>With the advancement of Industry 4.0 and intelligent manufacturing, there is an increasing demand for enhanced safety and reliability in equipment operation. As a precursor to mechanical failures, abnormal sounds are critical indicators, and their accurate detection plays a vital role in accident prevention and operational efficiency. To address the limitations of existing methods—such as heavy reliance on handcrafted acoustic features and model structures, insufficient detail representation, and poor cross-device robustness—this paper proposes an end-to-end detection framework based on Pre-trained Representation-driven and Multi-domain Feature Fusion (PReMFF). Specifically, we fine-tune a large-scale pre-trained model, Wav2vec 2.0, to extract generalized acoustic features. To further improve performance, we introduce two specialized modules: an adaptive frequency band enhancement module that highlights key frequency components, and a multi-scale dilated causal temporal modeling module that captures long-range dependencies in the time domain. Finally, the three-way features are gated and fused and jointly supervised by the classifier and loss function, achieving excellent performance of 94.29% and 88.88% pAUC on the DCASE 2020 TASK 2 dataset. The MIMII dataset is used to verify its ability to quickly adapt and robustly generalize under new equipment and complex noise, indicating that it provides an efficient and feasible solution for intelligent monitoring of industrial sites.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pre-trained representation-driven and multi-domain feature fusion method for anomalous sound detection

  • Jingwen Wei,
  • Hongjun Sun,
  • Chengyang Li,
  • Qian Wei,
  • Kaiwen Xing

摘要

With the advancement of Industry 4.0 and intelligent manufacturing, there is an increasing demand for enhanced safety and reliability in equipment operation. As a precursor to mechanical failures, abnormal sounds are critical indicators, and their accurate detection plays a vital role in accident prevention and operational efficiency. To address the limitations of existing methods—such as heavy reliance on handcrafted acoustic features and model structures, insufficient detail representation, and poor cross-device robustness—this paper proposes an end-to-end detection framework based on Pre-trained Representation-driven and Multi-domain Feature Fusion (PReMFF). Specifically, we fine-tune a large-scale pre-trained model, Wav2vec 2.0, to extract generalized acoustic features. To further improve performance, we introduce two specialized modules: an adaptive frequency band enhancement module that highlights key frequency components, and a multi-scale dilated causal temporal modeling module that captures long-range dependencies in the time domain. Finally, the three-way features are gated and fused and jointly supervised by the classifier and loss function, achieving excellent performance of 94.29% and 88.88% pAUC on the DCASE 2020 TASK 2 dataset. The MIMII dataset is used to verify its ability to quickly adapt and robustly generalize under new equipment and complex noise, indicating that it provides an efficient and feasible solution for intelligent monitoring of industrial sites.