错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing the Generalizability of Cancer Prognosis Models: Breast and Colon Cancer Case Studies

  • Wafaa Tizi,
  • Abdelaziz Berrado

摘要

Overview Machine learning provides tools to aid in decision-making in various domains. The benefit of these Machine learning techniques is also prevalent in the healthcare industry, especially in cancer prognosis. This paper aims to assess the generalizability of a set of models obtained from a variety of machine learning techniques for survival analysis, by testing the resulting prognosis models on different validation sets of unrelated populations. Methods Cox Proportional Hazards, Random Survival Forests, and Gradient Boosting for Survival Analysis were used to create breast and colon cancer prognosis models using SEER and SILU datasets respectively. The resulting models were tested on different datasets, namely: Duke and METABRIC breast cancer datasets, and TCGA and CPTAC colon cancer datasets. In order to perform this assessment, we start with a manual preprocessing step of feature matching that resulted in different subsets of the original datasets with a matching feature space. Models were created using each subset of SEER and SILU datasets, and their resulting C-index and Brier score were compared to the subsets of the external validation sets of Duke and METABRIC for breast cancer and TCGA and CPTAC for colon cancer. Results Different performance metrics give different results. Regardless of the technique used, the performance of the resulting model is not consistent as the validation set changes. The performance of the different models created with the SEER dataset dropped from an average C-index of 70% to 50% on the METABRIC validation set. The choice of the machine learning algorithm for survival analysis can affect the consistency of the resulting models