Breast cancer is a significant global health challenge, as it is the main cause of cancer-related deaths among women. It is a heterogeneous disease, manifesting in diverse clinical outcomes influenced by its intricate molecular characteristics. Accurate survival prediction for breast cancer patients is essential for guiding treatment strategies and improving their outcomes, which drives the exploration of various -omics datasets through advanced machine learning techniques. This study aims to enhance the accuracy of breast cancer survival prediction by introducing novel feature extraction methods combined with rigorous feature selection processes, leveraging multi-omics data for a more comprehensive analysis. Specifically, we evaluate the predictive power of various -omics datasets, including RNA, microRNA, and protein expression levels, DNA copy number variations, somatic mutation positions, and methylation levels. These datasets are used to create and compare single- and multi-omic models capable of predicting 5-year survival rates for breast cancer. Additionally, we assess the impact of different feature selection and data aggregation techniques on patient classification accuracy, as measured by the area under the ROC curve (AUC). Through the analysis of multi-omic data from 178 patients, we identified 35 key molecular predictors, which enabled the development of a robust 5-year survival prediction model, achieving an AUC of 0.89 in nested cross-validation. Our findings highlight the potential of integrating multi-omics data with advanced machine learning techniques to significantly improve breast cancer survival predictions. Moreover, we demonstrate that mRNA levels, gene copy number changes, and methylation levels are among the most effective predictors for this purpose.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Extraction and Selection of Multi-omic Features for the Breast Cancer Survival Prediction

  • Daria Kostka,
  • Wiktoria Płonka,
  • Roman Jaksik

摘要

Breast cancer is a significant global health challenge, as it is the main cause of cancer-related deaths among women. It is a heterogeneous disease, manifesting in diverse clinical outcomes influenced by its intricate molecular characteristics. Accurate survival prediction for breast cancer patients is essential for guiding treatment strategies and improving their outcomes, which drives the exploration of various -omics datasets through advanced machine learning techniques. This study aims to enhance the accuracy of breast cancer survival prediction by introducing novel feature extraction methods combined with rigorous feature selection processes, leveraging multi-omics data for a more comprehensive analysis. Specifically, we evaluate the predictive power of various -omics datasets, including RNA, microRNA, and protein expression levels, DNA copy number variations, somatic mutation positions, and methylation levels. These datasets are used to create and compare single- and multi-omic models capable of predicting 5-year survival rates for breast cancer. Additionally, we assess the impact of different feature selection and data aggregation techniques on patient classification accuracy, as measured by the area under the ROC curve (AUC). Through the analysis of multi-omic data from 178 patients, we identified 35 key molecular predictors, which enabled the development of a robust 5-year survival prediction model, achieving an AUC of 0.89 in nested cross-validation. Our findings highlight the potential of integrating multi-omics data with advanced machine learning techniques to significantly improve breast cancer survival predictions. Moreover, we demonstrate that mRNA levels, gene copy number changes, and methylation levels are among the most effective predictors for this purpose.