Acoustic scene classification is a task that has attracted the attention of numerous researchers in the audio classification research community. In this work, we evaluate the performance of handcrafted and non-handcrafted features when used individually, as well as the performance achieved by combining the outputs of classifiers individually fed with both these types of representation in the task of acoustic scene classification. For this purpose, we utilized the dataset from the DCASE 2016 challenge, which encompasses 15 different categories of acoustic scenes. To facilitate the comparison between handcrafted and non-handcrafted features, we converted the original audio signal into a visual representation in the time-frequency domain (i.e., spectrogram), as images provide a suitable format for feeding deep models for representation learning (i.e., non-handcrafted), and there are suitable texture operators described in the image processing literature to perform feature extraction in the handcrafted mode. Additionally, we assessed the impact of data augmentation strategies on the final classification results, given that the dataset used is relatively small, specially for the application of deep models. Experiments revealed the highest accuracy rate achieved using solely handcrafted features was 64.2%, while it reached 91.2% using non-handcrafted features. The best overall accuracy rate was attained when we combined classifiers created using non-handcrafted features with classifiers created using handcrafted features, resulting in an accuracy rate of 92.5%. This corroborates the hypothesis raised in this work, that there is complementarity between both types of representation in the task investigated.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hancrafted Vs. Non-handcrafted Features for Acoustic Scene Classification

  • João Vitor S. Castanho,
  • Yandre M. G. Costa

摘要

Acoustic scene classification is a task that has attracted the attention of numerous researchers in the audio classification research community. In this work, we evaluate the performance of handcrafted and non-handcrafted features when used individually, as well as the performance achieved by combining the outputs of classifiers individually fed with both these types of representation in the task of acoustic scene classification. For this purpose, we utilized the dataset from the DCASE 2016 challenge, which encompasses 15 different categories of acoustic scenes. To facilitate the comparison between handcrafted and non-handcrafted features, we converted the original audio signal into a visual representation in the time-frequency domain (i.e., spectrogram), as images provide a suitable format for feeding deep models for representation learning (i.e., non-handcrafted), and there are suitable texture operators described in the image processing literature to perform feature extraction in the handcrafted mode. Additionally, we assessed the impact of data augmentation strategies on the final classification results, given that the dataset used is relatively small, specially for the application of deep models. Experiments revealed the highest accuracy rate achieved using solely handcrafted features was 64.2%, while it reached 91.2% using non-handcrafted features. The best overall accuracy rate was attained when we combined classifiers created using non-handcrafted features with classifiers created using handcrafted features, resulting in an accuracy rate of 92.5%. This corroborates the hypothesis raised in this work, that there is complementarity between both types of representation in the task investigated.