A feature selection method based on multi-scale information fusion
摘要
With the continuous growth of data scale and dimensionality, effective feature selection from massive datasets has become crucial in data mining and machine learning. As an effective mathematical tool for processing uncertain data, rough set theory provides theoretical support for feature selection. However, when dealing with continuous data, traditional rough set methods mainly rely on single-scale discretization strategies and single-scale indicators to evaluate attribute importance, which can easily lead to information loss and reduced reduction accuracy. To overcome these limitations, this paper proposes a multi-scale information-fusion feature selection algorithm. Continuous data are discretized at multiple scales, with conditional entropy and dependency calculated for each scale. Then, a joint evaluation index is designed, and a dynamic weight adaptive-adjustment mechanism is introduced to achieve effective reduction of continuous data. Experiments on UCI datasets demonstrate that the proposed method outperforms existing algorithms in feature reduction performance.