Particularly in biomedical applications, feature selection plays a critical role in enhancing the interpretability and efficacy of machine learning models. This work examines the performance of Explainable AI (XAI), Information Gain (IG), and Principal Component Analysis (PCA) methods on an imbalanced dataset pertaining to stroke prediction. Data from 4,603 patient records, including 362 instances of stroke from the National Health and Nutrition Examination Survey are used in this investigation. Methodologically, IG is used for feature ranking, PCA is used to reduce dimensionality and XAI techniques are used to improve model transparency. The chosen features are used to assess the performance of several machine learning models, including Random Forest, Support Vector Machine, k-Nearest Neighbours, and Logistic Regression, in terms of classification. Our experimental results show that the combined PCA-IG approach significantly enhances classification accuracy, achieving 91.75%. Furthermore, LIME-based feature selection outperformed in precision, recall, and F1 score, with the highest accuracy at 91.86%. LIME discovered nine positive impact features, highlighting the top contributors in the dataset. We also applied the same feature selection technique to datasets from other domains. These findings highlight the robustness of using PCA-IG and XAI approaches separately to create reliable and understandable machine learning models for healthcare and other applications. By offering insights into the optimal use of PCA, IG, and XAI to enhance the accuracy and practicality of machine learning models in healthcare and other domains, this paper advances the field of feature selection across all areas of data analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable AI in Feature Selection: Improving Classification Performance on Imbalanced Datasets

  • Shahriar Siddique Ayon,
  • Muhammad Ebrahim Hossain,
  • Md Saef Ullah Miah,
  • M. Mostafizur Rahman,
  • Mufti Mahmud

摘要

Particularly in biomedical applications, feature selection plays a critical role in enhancing the interpretability and efficacy of machine learning models. This work examines the performance of Explainable AI (XAI), Information Gain (IG), and Principal Component Analysis (PCA) methods on an imbalanced dataset pertaining to stroke prediction. Data from 4,603 patient records, including 362 instances of stroke from the National Health and Nutrition Examination Survey are used in this investigation. Methodologically, IG is used for feature ranking, PCA is used to reduce dimensionality and XAI techniques are used to improve model transparency. The chosen features are used to assess the performance of several machine learning models, including Random Forest, Support Vector Machine, k-Nearest Neighbours, and Logistic Regression, in terms of classification. Our experimental results show that the combined PCA-IG approach significantly enhances classification accuracy, achieving 91.75%. Furthermore, LIME-based feature selection outperformed in precision, recall, and F1 score, with the highest accuracy at 91.86%. LIME discovered nine positive impact features, highlighting the top contributors in the dataset. We also applied the same feature selection technique to datasets from other domains. These findings highlight the robustness of using PCA-IG and XAI approaches separately to create reliable and understandable machine learning models for healthcare and other applications. By offering insights into the optimal use of PCA, IG, and XAI to enhance the accuracy and practicality of machine learning models in healthcare and other domains, this paper advances the field of feature selection across all areas of data analysis.