DTST-MFN: Enhancing Anomalous Sound Detection with Deep Temporal Features and Efficient Attention
摘要
In the wake of deep learning's rapid development, the realm of machine anomalous sound detection has notably advanced through self-supervised learning techniques. Despite significant progress, existing models display deficiencies in adequately utilizing deep temporal audio features. To address this shortfall, this study proposes Deep-TgramNet, a novel feature extraction framework that, by markedly increasing the convolutional depth and widening the receptive field, achieves the effective utilization of temporal features neglected by existing models, thereby significantly enhancing the model's ability to discern between normal and abnormal samples. In addition, we developed the DTST-MFN framework by integrating the Deep-TgramNet, STgram-MFN, and efficient attention mechanisms. DTST-MFN not only fully extracts temporal features but also focuses on key features in the temporal and frequency domains that are beneficial for anomalous sound detection, and adaptively adjusts the weights between channels. Experiments on the DCASE2020 Task 2 dataset demonstrate DTST-MFN's superiority in anomaly sound detection, achieving an average AUC of 94.06% and pAUC of 88.79%.