The diagnosis of cancer is a difficult task due to symptoms being confused as benign when the time to diagnose is critical in the survival of patients. The situation is compounded by rare cancer types which general practitioners and patients alike can only assume may be possible when all other explanations have been exhausted. Multilabel classification of cancer subtypes in health survey data is challenged by class imbalance and erroneous data. Many prior studies focus on the classification of single cancer types, and health information for avoiding or surviving cancers is typically isolated rather than generalized. Knowledge of chronic illness risk factors can aid healthier lifestyles and assist the recognition of signs requiring prompt medical attention. By transforming survey keywords into linear variables, we merge annual Behavioural and Risk Factor Surveillance System (BRFSS) health surveys, to boost rare cancer subtype samples of data, and enable optimal Synthetic Minority Oversampling Technique (SMOTE) performance. Our iterative method of risk factor and symptom feature selection viz. Alcohol (A), Diabetes (D), Gender + Weight (GW), Gender + Injury + Fatigue (GIF), and Smoking (S) deliver an explainable predictive machine learning (ML) model for many cancer subtypes and cancer occurrences.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Predictive Accuracy of Multilabel Classification for Cancer Subtypes in Imbalanced Datasets

  • M. Jaworsky,
  • X. Tao,
  • T. Shaik,
  • J. Yong,
  • L. Pan,
  • J. Zhang,
  • S. R. Pokhrel

摘要

The diagnosis of cancer is a difficult task due to symptoms being confused as benign when the time to diagnose is critical in the survival of patients. The situation is compounded by rare cancer types which general practitioners and patients alike can only assume may be possible when all other explanations have been exhausted. Multilabel classification of cancer subtypes in health survey data is challenged by class imbalance and erroneous data. Many prior studies focus on the classification of single cancer types, and health information for avoiding or surviving cancers is typically isolated rather than generalized. Knowledge of chronic illness risk factors can aid healthier lifestyles and assist the recognition of signs requiring prompt medical attention. By transforming survey keywords into linear variables, we merge annual Behavioural and Risk Factor Surveillance System (BRFSS) health surveys, to boost rare cancer subtype samples of data, and enable optimal Synthetic Minority Oversampling Technique (SMOTE) performance. Our iterative method of risk factor and symptom feature selection viz. Alcohol (A), Diabetes (D), Gender + Weight (GW), Gender + Injury + Fatigue (GIF), and Smoking (S) deliver an explainable predictive machine learning (ML) model for many cancer subtypes and cancer occurrences.