<p>This work presents <Emphasis FontCategory="NonProportional">AudioSet-Tools</Emphasis>, a modular and extensible Python framework designed to streamline the creation of task-specific datasets derived from Google AudioSet. Despite its extensive coverage, AudioSet suffers from weak labeling, class imbalance, and a loosely structured taxonomy, which hinder its applicability in <i>machine listening</i> workflows. <Emphasis FontCategory="NonProportional">AudioSet-Tools</Emphasis> addresses these issues through configurable taxonomy-consistent label filtering and class rebalancing strategies. The framework includes automated routines for data download and transformation, enabling reproducible and semantically consistent dataset generation for pre-training and downstream fine-tuning of deep learning models. While domain-agnostic, we showcase its versatility through <Emphasis FontCategory="NonProportional">AudioSet-EV</Emphasis>, a curated subset focused on emergency vehicle siren recognition — a socially relevant and technically challenging use case that highlights structural and semantic gaps in the AudioSet taxonomy. We further provide an extensive comparative benchmark of <Emphasis FontCategory="NonProportional">AudioSet-EV</Emphasis> against state-of-the-art emergency vehicle corpora. All source code and datasets are openly released on <a href="https://github.com/StefanoGiacomelli/audioset-tools/">GitHub</a> and <a href="https://zenodo.org/records/14882314">Zenodo</a>, fostering transparency and reproducibility in real-world audio signal processing research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AudioSet-tools: a Python framework for taxonomy-aware AudioSet curation and reproducible audio research

  • Stefano Giacomelli,
  • Marco Giordano,
  • Claudia Rinaldi,
  • Fabio Graziosi

摘要

This work presents AudioSet-Tools, a modular and extensible Python framework designed to streamline the creation of task-specific datasets derived from Google AudioSet. Despite its extensive coverage, AudioSet suffers from weak labeling, class imbalance, and a loosely structured taxonomy, which hinder its applicability in machine listening workflows. AudioSet-Tools addresses these issues through configurable taxonomy-consistent label filtering and class rebalancing strategies. The framework includes automated routines for data download and transformation, enabling reproducible and semantically consistent dataset generation for pre-training and downstream fine-tuning of deep learning models. While domain-agnostic, we showcase its versatility through AudioSet-EV, a curated subset focused on emergency vehicle siren recognition — a socially relevant and technically challenging use case that highlights structural and semantic gaps in the AudioSet taxonomy. We further provide an extensive comparative benchmark of AudioSet-EV against state-of-the-art emergency vehicle corpora. All source code and datasets are openly released on GitHub and Zenodo, fostering transparency and reproducibility in real-world audio signal processing research.