Self-supervised methods have achieved state-of-the-art performance in Anomalous Sound Detection (ASD) task, which exploits manually annotated meta information (i.e., section IDs and attributes) as labels. However, these methods are not practical in the real world that the scenarios of domain shifts represented by section IDs are complex and diverse, which usually do not emerge entirely in the training data. Diffusion models have achieved outstanding performance in various tasks related to generative modeling for their powerful generative and generalization capabilities. In this paper, we propose the ASD-Diff, an unsupervised Anomalous Sound Detection method with masked Diffusion model, which adopts a one-step denoising approximation in the backward diffusion process that makes the inference faster than the general diffusion methods and is proven to improve the ASD performance. We aim to leverage diffusion models to distinguish between normal and abnormal sounds by learning the distribution of normal acoustic features. We conducted the experiments on the Task 2 dataset of the DCASE 2022 Challenge, which is the first ASD dataset for domain generalization techniques. Results show that the ASD-Diff outperforms six compared methods on the whole.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ASD-Diff: Unsupervised Anomalous Sound Detection with Masked Diffusion Model

  • Xin Fan,
  • Wenjie Fang,
  • Ying Hu

摘要

Self-supervised methods have achieved state-of-the-art performance in Anomalous Sound Detection (ASD) task, which exploits manually annotated meta information (i.e., section IDs and attributes) as labels. However, these methods are not practical in the real world that the scenarios of domain shifts represented by section IDs are complex and diverse, which usually do not emerge entirely in the training data. Diffusion models have achieved outstanding performance in various tasks related to generative modeling for their powerful generative and generalization capabilities. In this paper, we propose the ASD-Diff, an unsupervised Anomalous Sound Detection method with masked Diffusion model, which adopts a one-step denoising approximation in the backward diffusion process that makes the inference faster than the general diffusion methods and is proven to improve the ASD performance. We aim to leverage diffusion models to distinguish between normal and abnormal sounds by learning the distribution of normal acoustic features. We conducted the experiments on the Task 2 dataset of the DCASE 2022 Challenge, which is the first ASD dataset for domain generalization techniques. Results show that the ASD-Diff outperforms six compared methods on the whole.