A high-dimensional feature selection method based on feature interaction clustering and integer-encoded TLBO
摘要
High-dimensional feature selection faces significant challenges in identifying optimal feature subsets, especially with low-sample-size data. Most traditional methods that assess redundancy through pairwise feature correlation risk eliminating features with complementary classification effects, while existing evolutionary algorithms often suffer from suboptimal global exploration and slow convergence rates. To address these challenges, we propose a novel Adaptive Redundancy Clustering-TLBO (ARC-TLBO) algorithm. A new metric, label-specific information gain, is first introduced to measure feature redundancy from the perspective of feature interaction. Building upon this metric, an adaptive clustering method is proposed to group redundant features, thereby reducing the search space for evolutionary algorithms. Furthermore, the teaching-learning-based optimization (TLBO) algorithm is improved to identify the most representative features from each group. Based on positive and reverse learning, the teaching and learning operators are introduced to accelerate the search for the optimal feature subset. An elite retention mechanism preserves high-quality solutions, improving stability and preventing overfitting. Comprehensive experimental results on 18 high-dimensional low-sample-size datasets show that ARC-TLBO achieves higher accuracy, faster convergence, and more compact feature subsets, making it highly effective for feature selection tasks.