A Stacking Ensemble Machine Learning Approach Based on Urinary Biomarkers to Diagnose Pancreatic Cancer
摘要
Pancreatic cancer (PC) ranks as one of the serious and aggressive types of cancer, with a higher mortality rate due to a lack of accurate diagnosis in the initial phases. Therefore, well-timed detection of pancreatic cancer is critical for improving its prognosis. Based on medical data, machine learning (ML) algorithms have been proposed as potential techniques to enhance cancer detection. The present research aims to develop a novel stacking ensemble model using a comprehensive pre-processing method and various machine learning techniques to diagnose pancreatic cancer using urinary and CA 19–9 biomarkers. Specifically, we utilized an advanced pre-processing pipeline, which included managing missing values, encoding, Boruta feature selection, and oversampling by K-Means Synthetic Minority Over-Sampling Technique (K-Means SMOTE) to enhance data quality. Furthermore, we contrasted the performance of the proposed ensemble learning algorithm with that of eight individual and traditional machine learning algorithms, employing a stacking ensemble with Random Forest as a meta-learner, trained via K-Fold cross-validation to optimize predictive performance. According to the study's findings, the proposed stacking ensemble model outperforms other individual models for early diagnosis of pancreatic cancer, achieving an accuracy of 97.52% and an area under the curve (AUC) of 99.19% on the provided features. Next, we sorted perioperative variables based on the Local Interpretable Model-agnostic Explanations (LIME) value to identify the most significant features. This robust feature selection, combined with sophisticated machine learning, can improve pancreatic cancer detection and patient outcome management.