Accurate prediction of breast cancer recurrence is vital for guiding patient treatment strategies and improving long-term outcomes. This study proposes a machine learning-based approach that utilizes the METABRIC dataset, incorporating both genetic and clinical features to assess the likelihood of cancer relapse. Our methodology includes advanced preprocessing techniques to handle missing data, employing a combination of Random Forest imputation and statistical methods to ensure data integrity. Following data preparation, multiple classification models including Random Forest, Decision Tree, XGBoost, Logistic Regression, and a Dense Neural Network were trained to classify patients by recurrence status. Model performance was thoroughly evaluated using key metrics, with AUC–ROC analysis highlighting predictive accuracy of thse different algorithms. The results indicate that machine learning-driven approaches, when combined with robust data preprocessing, can significantly improve the reliability of breast cancer recurrence predictions. This framework not only offers insights for personalized patient care but also provides a scalable foundation for future research aimed at enhancing prognostic tools in oncology. The modular nature of the system allows for adaptation and optimization, positioning this work as a valuable contribution to the field of cancer prognosis and personalized medicine.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Breast Cancer Prognosis: Machine Learning Models for Recurrence Prediction

  • Mohamed Salah Amine Benouar,
  • Seifeddine Bouallegue

摘要

Accurate prediction of breast cancer recurrence is vital for guiding patient treatment strategies and improving long-term outcomes. This study proposes a machine learning-based approach that utilizes the METABRIC dataset, incorporating both genetic and clinical features to assess the likelihood of cancer relapse. Our methodology includes advanced preprocessing techniques to handle missing data, employing a combination of Random Forest imputation and statistical methods to ensure data integrity. Following data preparation, multiple classification models including Random Forest, Decision Tree, XGBoost, Logistic Regression, and a Dense Neural Network were trained to classify patients by recurrence status. Model performance was thoroughly evaluated using key metrics, with AUC–ROC analysis highlighting predictive accuracy of thse different algorithms. The results indicate that machine learning-driven approaches, when combined with robust data preprocessing, can significantly improve the reliability of breast cancer recurrence predictions. This framework not only offers insights for personalized patient care but also provides a scalable foundation for future research aimed at enhancing prognostic tools in oncology. The modular nature of the system allows for adaptation and optimization, positioning this work as a valuable contribution to the field of cancer prognosis and personalized medicine.