Synthetic Data for Feature Selection
摘要
This study aims to perform a comparative analysis of popular feature selection algorithms in the regime of imbalanced data using synthetic datasets. The use of synthetic data has several advantages over the traditional, real-life data including the direct evaluation of the algorithms without the classifier and metrics bias. As a result, we obtain a more accurate evaluation of the algorithms than existing approaches. The results—based on eight different algorithms and eight different datasets—show that sequential forward selection and entropy-based selection achieve the highest precision in the regime of imbalanced data.