Responsible Music Genre Classification Using Interpretable Model-Agnostic Visual Explainers
摘要
The responsible artificial intelligence (AI) paradigm requires machine learning (ML) and AI engineers to ensure transparent and interpretable intelligent models across domains. This requirement becomes even more complex and crucial when dealing with feature learning using sound data, as deriving model explainers in this context is a more detailed and sophisticated process. In this work, we demonstrate a responsible approach to AI modeling and leverage three explainable artificial intelligence (XAI) tools to derive and establish explainers for the music genre classification model-agonistically. Explain like am 5 (ELi5), Shapley Additive exPlanations (SHAP), and Local Interpretable Model-agnostic Explanations (LIME) were particularly used to interpret and simplify the model’s decision-making processes regardless of their discrete mathematical foundations. This transparency provides a deeper understanding of sound feature patterns and characteristics that inform genre classes for the models through graphical and qualitative feature contributions to explain the model’s justification. We developed convolutions neural networks (CNNs) with and without cross-validation and a vision transformer (ViT) approach utilizing MobileNets. Ultimately, the CNN with cross-validation demonstrated superior performance, achieving 80% accuracy on the test set and 84% accuracy on the validation set. This work advances the border of music intelligence research and promotes the broader cause of responsible AI to ensure that complex models remain comprehensible and accountable.