A Unified Machine Learning Framework for Multi-subtype Tumour Classification Across Diverse Datasets
摘要
Machine learning models are increasingly employed in the classification of digitized medical tissue images, including for identifying cancer types and subtypes. Most models focus on a target tumor type, using datasets from a single source, which limits generalization. To overcome this, we formulated unified frameworks and applied them to bone, colon, and prostate cancer datasets from varied origins. Rigorous testing concluded that framework 1 achieved an overall accuracy of 88.48%, while framework 2, with classification corrections, achieved an enhanced overall accuracy of 90.28%. The area under the curve (AUC) values exceeded 0.97, with average specificity and sensitivity surpassing 0.98 and 0.99 across frameworks for individual classes. Additionally, when Framework 2 was applied to an unseen breast cancer dataset, it demonstrated a notable accuracy of 88.8% for normal vs. tumor tile classification, indicating its efficacy on unseen datasets. The results reflect the potential of engineered feature extraction in predictive models.