Setting Importance of Features Through Means and Majority of Outcomes of Machine Learning Algorithms: An Empirical Analysis
摘要
This study attempts to set the importance of features that influence the target variable. A set of algorithms, namely, Random Forest (RF), Minimum Redundancy Maximal Relevancy Ensemble (mRMRE), Boruta, and Linear Regression (LR) are employed to establish the relationships among the features. Then, a unified score is generated for each feature using the outputs of all algorithms through min–max normalization and referring geometric, arithmetic, and harmonic means. Finally, an aggregate rank is suggested for each feature by a heuristic convention on the said unified scores. The proposed methodology is tested by finding the importance of the attributes concerning the sales for a garment dataset. The aim is to assign universal importance to each feature that highly kindles the sales of dresses at an aggregate level. Mention that preliminary filtering and augmentation of the data have been executed to prepare the case study. It is evident from the results that the importance of features is graded adequately and clearly separable through the proposed methodology based on a unified score and convention of aggregate ranking by processing the outcomes of multiple machine learning algorithms.