<p>Music genre classification is a fundamental task in multimedia information retrieval, with broad applications in music recommendation, content indexing, and large-scale audio analysis. Despite significant progress, most existing approaches rely on fully supervised learning and require large amounts of labeled data, which limits their applicability in low-resource scenarios. Recent audio–language models enable zero-shot audio classification by aligning audio signals with natural language descriptions, providing a promising direction for label-efficient music understanding. However, their performance is often sensitive to prompt design and degrades notably when only a few labeled examples are available. In this work, we propose a unified zero-to-few shot framework for music genre classification based on audio–language models. The framework integrates prototype calibration, which adaptively fuses text-based and audio-based class representations, with a parameter-efficient embedding-space prompt learning strategy. The proposed approach keeps the backbone model frozen and introduces only a small number of trainable parameters, resulting in low computational overhead and stable optimization. Extensive experiments on GTZAN, FMA-small, and FMA-medium show that the proposed method achieves competitive and often improved performance compared with standard zero-shot prompting, recent representative baselines, and lightweight few-shot competitors under the evaluated settings. The gains are most evident in low-shot adaptation and parameter efficiency, while the FMA-medium results also indicate that class imbalance remains a challenging factor, especially in terms of macro-F1.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-to-few shot music genre classification via prototype calibration and prompt learning in audio–language models

  • Jiahong Shen,
  • Guangrun Xiao

摘要

Music genre classification is a fundamental task in multimedia information retrieval, with broad applications in music recommendation, content indexing, and large-scale audio analysis. Despite significant progress, most existing approaches rely on fully supervised learning and require large amounts of labeled data, which limits their applicability in low-resource scenarios. Recent audio–language models enable zero-shot audio classification by aligning audio signals with natural language descriptions, providing a promising direction for label-efficient music understanding. However, their performance is often sensitive to prompt design and degrades notably when only a few labeled examples are available. In this work, we propose a unified zero-to-few shot framework for music genre classification based on audio–language models. The framework integrates prototype calibration, which adaptively fuses text-based and audio-based class representations, with a parameter-efficient embedding-space prompt learning strategy. The proposed approach keeps the backbone model frozen and introduces only a small number of trainable parameters, resulting in low computational overhead and stable optimization. Extensive experiments on GTZAN, FMA-small, and FMA-medium show that the proposed method achieves competitive and often improved performance compared with standard zero-shot prompting, recent representative baselines, and lightweight few-shot competitors under the evaluated settings. The gains are most evident in low-shot adaptation and parameter efficiency, while the FMA-medium results also indicate that class imbalance remains a challenging factor, especially in terms of macro-F1.