Machine learning models are increasingly employed in the classification of digitized medical tissue images, including for identifying cancer types and subtypes. Most models focus on a target tumor type, using datasets from a single source, which limits generalization. To overcome this, we formulated unified frameworks and applied them to bone, colon, and prostate cancer datasets from varied origins. Rigorous testing concluded that framework 1 achieved an overall accuracy of 88.48%, while framework 2, with classification corrections, achieved an enhanced overall accuracy of 90.28%. The area under the curve (AUC) values exceeded 0.97, with average specificity and sensitivity surpassing 0.98 and 0.99 across frameworks for individual classes. Additionally, when Framework 2 was applied to an unseen breast cancer dataset, it demonstrated a notable accuracy of 88.8% for normal vs. tumor tile classification, indicating its efficacy on unseen datasets. The results reflect the potential of engineered feature extraction in predictive models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Unified Machine Learning Framework for Multi-subtype Tumour Classification Across Diverse Datasets

  • Ankur Yadav,
  • Ovidiu Daescu

摘要

Machine learning models are increasingly employed in the classification of digitized medical tissue images, including for identifying cancer types and subtypes. Most models focus on a target tumor type, using datasets from a single source, which limits generalization. To overcome this, we formulated unified frameworks and applied them to bone, colon, and prostate cancer datasets from varied origins. Rigorous testing concluded that framework 1 achieved an overall accuracy of 88.48%, while framework 2, with classification corrections, achieved an enhanced overall accuracy of 90.28%. The area under the curve (AUC) values exceeded 0.97, with average specificity and sensitivity surpassing 0.98 and 0.99 across frameworks for individual classes. Additionally, when Framework 2 was applied to an unseen breast cancer dataset, it demonstrated a notable accuracy of 88.8% for normal vs. tumor tile classification, indicating its efficacy on unseen datasets. The results reflect the potential of engineered feature extraction in predictive models.