Machine learning approaches for predicting breast cancer recurrence using clinical and histopathological data
摘要
Breast cancer remains the most common malignancy among women worldwide, with recurrence representing a major clinical challenge. Although significant progress has been made in early detection and treatment, recurrence affects up to 40% of patients in Brazil, influencing survival outcomes and therapeutic decisions. In this context, Machine Learning offers valuable potential for enhancing recurrence prediction by enabling data-driven risk assessment and personalized patient care. In the present study, clinical and histopathological information was extracted from unstructured medical records of breast cancer patients. A clustering technique (K-Means) was applied to identify patient subgroups with varying tumor aggressiveness profiles. Survival outcomes were further analyzed using the Cox proportional hazards model. Two distinct subgroups were identified: for less aggressive tumors, Quadratic Discriminant Analysis achieved a remarkably high recall of 0.9872, while for more aggressive tumors, Random Forest provided the most favorable trade-off between recall (0.7296) and precision (0.6811). Future research should explore validation across multiple institutions, incorporate molecular biomarkers, and leverage deep learning approaches to enhance predictive performance.