<p>Feature selection remains a central topic of research across various application domains. It is widely employed to reduce dimensionality, eliminate irrelevant data, enhance learning accuracy, and improve the interpretability of results. Among the various feature selection techniques, wrapper methods are widely recognized for their strong performance, albeit at the cost of high computational expense. In this paper, we implement a general framework for the ensemble of multiple feature selection methods based on bootstrap-induced diversity. This framework offers a streamlined methodology for enhancing both the predictive accuracy and stability of wrapper-based feature selection methods. Experimental results, derived from five simulated datasets and ten real-world datasets, underscore the effectiveness of regression-based wrappers in supervised learning problems. Notably, among the real datasets, partial least squares and logworth scoring methods demonstrate superior accuracy and stability overall. Furthermore, the degree of collinearity in the input data emerges as a decisive factor in determining the optimal approach between these two methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Within Importance Score Aggregation for Wrapper-Based Feature Selection and Its Stability

  • Reem Salman,
  • Ayman Alzaatreh,
  • Hana Sulieman

摘要

Feature selection remains a central topic of research across various application domains. It is widely employed to reduce dimensionality, eliminate irrelevant data, enhance learning accuracy, and improve the interpretability of results. Among the various feature selection techniques, wrapper methods are widely recognized for their strong performance, albeit at the cost of high computational expense. In this paper, we implement a general framework for the ensemble of multiple feature selection methods based on bootstrap-induced diversity. This framework offers a streamlined methodology for enhancing both the predictive accuracy and stability of wrapper-based feature selection methods. Experimental results, derived from five simulated datasets and ten real-world datasets, underscore the effectiveness of regression-based wrappers in supervised learning problems. Notably, among the real datasets, partial least squares and logworth scoring methods demonstrate superior accuracy and stability overall. Furthermore, the degree of collinearity in the input data emerges as a decisive factor in determining the optimal approach between these two methods.