Comparative Analysis of Data Preprocessing Methods in Machine Learning for Breast Cancer Classification
摘要
Breast cancer, the most diagnosed cancer according to the American Cancer Society, impacts about 13% of women and has a rising incidence rate. With a 10-year survival rate of 84%, early detection is essential, especially in cases of dense breast tissue where 41% go undetected. This paper explores various feature selection and extraction techniques in machine learning for breast cancer classification, including correlation-based selection, recursive elimination, linear discriminant analysis, principal component analysis, and their combinations. The experimental evaluation with the well-known Wisconsin Diagnostic Breast Cancer dataset shows that the linear discriminant analysis feature transformation provides the best overall performance for machine learning algorithmic classifications regarding Accuracy, Precision, Recall, and F1-Score.