On the utility of ordered incremental attribute learning-based variance and mRMR techniques
摘要
From the major issues of machine learning and data mining is feature selection. This consists in picking only the most significant features to further emphasize the learning search. One relaxed category of feature selection is feature ordering that orders the features according to their relevance. However, numerous feature selection techniques are still rarely used in unsupervised tasks as well as for incrementally clustering continuously emerging data streams with added mixed features. This paper provides a better insight into one particular machine learning strategy: the ordered incremental attribute learning based on K-prototypes. This means to investigate the impact of feature ordering preprocessing technique before modeling data streams with added mixed features. At the outset, our proposal selects and ranks the most relevant features from the new emerging ones. Then, it dynamically injects them in our proposed incremental K-prototypes algorithm. To do so, two feature ordering techniques have been applied. One is the mRMR feature selection technique, designed to determine the smallest subset of pertinent features with standing for minimum redundancy and maximum relevance. Two is sorting the new added features with respect to their variances. The comparative study between the different orders, batch K-prototypes and non-ordering incremental K-prototypes method reveals that our proposals are skilled and efficient according to the run time, the sum of squared error and the Davies–Bouldin index evaluation measures.