Improving the Reliability of Tree-Based Feature Importance via Consensus Signals
摘要
Feature (variable) selection methods are used to detect the most important features (variables) within high-dimensional data. Tree-based models, such as Random Forests, are often exploited for this purpose, as they provide a built-in mechanism to quantify feature importance. However, the stochastic sampling strategies used in these models can lead to unstable feature importance rankings, particularly when the number of trees is low. In our study, we investigate the extent to which these unstable feature rankings can be consolidated through rank aggregation and consensus signal techniques. We propose to compute consensus values from multiple feature selection runs, where each run generates a ranked list of features. We have evaluated our approach while varying a spectrum of hyperparameters such as the number of trees and the number of features available for splitting a node. Our results suggest that consensus ranks provide a more accurate and robust selection of features compared to single-run feature selection procedures. The proposed approach is especially relevant for biomarker discovery, as it can improve the accuracy and reliability of feature selection, leading to the identification of the most informative and relevant biomarkers associated with a particular disease or condition. By consolidating the results from multiple feature selection runs, our approach may help to overcome problems associated with noisy or complex data because it can provide more robust and accurate estimates of the feature importance rankings.