Cancer is one of the deadly diseases characterized by abnormal cell division, destructive to mankind. It has been one of the leading cause of deaths for decades worldwide. In this context, this study is designed to embrace the proficiency of ML algorithms in detecting the risk of cancer occurrence based on medical and lifestyle habits amongst individuals. Specifically, the current study is designed to satisfy three objectives namely, (1) identifying cancer-causing data attributes, (2) employing ML models to decipher risk of cancer, (3) improving performance of ML models to predict cancer instances. In this perspective, cancer dataset is collected from Kaggle repository and subjected to attribute selection approaches followed by classification algorithms. Experimentations revealed that eight prominent features induced risk of cancer, as identified by Boruta algorithm. Subsequently, nine ML models were employed on data to detect risk of cancer. The hybrid ensemble model, XRandGrad achieved enhanced cancer predictive performance with an accuracy score of 99.19%. Finally, the model was validated to ascertain its performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Machine Learning Ensemble Algorithm to Predict Occurrence of Cancer

  • Kalyan Nagaraj,
  • H. S. Prashanth,
  • Amulyashree Sridhar

摘要

Cancer is one of the deadly diseases characterized by abnormal cell division, destructive to mankind. It has been one of the leading cause of deaths for decades worldwide. In this context, this study is designed to embrace the proficiency of ML algorithms in detecting the risk of cancer occurrence based on medical and lifestyle habits amongst individuals. Specifically, the current study is designed to satisfy three objectives namely, (1) identifying cancer-causing data attributes, (2) employing ML models to decipher risk of cancer, (3) improving performance of ML models to predict cancer instances. In this perspective, cancer dataset is collected from Kaggle repository and subjected to attribute selection approaches followed by classification algorithms. Experimentations revealed that eight prominent features induced risk of cancer, as identified by Boruta algorithm. Subsequently, nine ML models were employed on data to detect risk of cancer. The hybrid ensemble model, XRandGrad achieved enhanced cancer predictive performance with an accuracy score of 99.19%. Finally, the model was validated to ascertain its performance.