错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Feature Selection Methods Based on Feature Clustering and Information Theory

  • Wenhui Wang,
  • Changyin Zhou

摘要

In order to identify the most representative subset of features in high-dimensional data, a feature selection algorithm (AP-MSU) based on feature clustering and information theory is proposed. The algorithm introduces the AP clustering algorithm and multivariate symmetric uncertainty (MSU) based on the filtering feature selection algorithm’s preliminary screening of relevant features, better demonstrating the interactions between multiple feature variables and their interactions with target variables. The features are evaluated sequentially by an MSU-based feature quality metric, which considers both redundancy and interaction among the candidate features in the selected feature set, and removes the redundant features by assessing the ability of the features to provide effective categorization information with a small amount of computation. The experimental results show that the AP-MSU feature selection algorithm can effectively select a good feature set on binary and multi-classified gene expression datasets, and has good classification effect on different classifiers. In addition, the classification accuracy can be improved by the algorithm obtained a lower dimensional subset of features.