Machine Learning for Diagnosis of Pancreatic Ductal Adenocarcinoma Using Urine Samples
摘要
Pancreatic ductal adenocarcinoma (PDAC), is the commonplace kind of pancreatic adenocarcinoma or PC. It is one of the most deadly and aggressive cancers with limited early-stage detection methods. It is one of the most dangerous cancers because of poor 5-year survivability of less than 10% and symptoms which surface by the time the later stages of disease are reached. Thus, there is a need for advanced diagnostic methods for PDAC and there has been an increased focus on discovery of methods. Many biomarkers have been discovered for detection of PDAC and these have been found in places such as in muscles, tissues and other parts of the body. The purpose of this paper was to train a model that combines biomarkers and machine learning algorithms to efficiently classify patients as either suffering from PDAC or not. The raw data of PDAC patients was taken from Kaggle. Machine learning algorithms, including forests (RF), and K-nearest neighbors (KNN), were trained using the selected biomarkers. These were tested using different techniques to assess their performance. The combination of such biomarkers, along with machine learning algorithms, achieved a high classification accuracy for PDAC diagnosis. The Random Forest classification algorithm achieved the highest accuracy of 97.97% with recall score 1. The developed model not only provides efficient and non-invasive PDAC classification but also offers potential insights into the underlying biological processes associated with the disease. The identified urine biomarkers hold promise as potential diagnostic and prognostic tools for PDAC. Furthermore, the machine learning-based classification model can be integrated into existing clinical workflows, aiding in early-stage detection and improving patient outcomes. The model demonstrates promising results in accurately identifying PDAC-positive patients. Further validation studies and clinical trials are warranted to assess the robustness and generalizability of the proposed model in a larger patient population.