Improved Gastric Cancer Diagnosis with Machine Learning Technique: Addressing Imbalanced Data Distribution
摘要
The use of computer-aided detection (CAD) system utilizing machine learning (ML) techniques can reduce the mortality rate of gastric cancer (GC) patients. Nevertheless, the presence of class imbalance within datasets has the potential to introduce bias into the predictions made by these systems which can lead to the misdiagnosis for a higher risk individual. This study addressed this problem by dividing the acquired 145,787 screening records from the NHS Liverpool hospital into imbalanced and balanced datasets by implementing the stratified sampling method. Four frequently utilized ML techniques, namely Support Vector Machine (SVM), Multilayer Perceptron (MLP), Naive Bayes (NB), and logistic regression (LR), were trained and evaluated on these two datasets across five distinct test circumstances. The obtained findings indicate that the utilization of a balanced dataset has the potential to enhance the accuracy of stomach cancer diagnosis. The dataset that was not balanced had a higher sensitivity for Class 0, indicating a lower number of missing negative cases. Conversely, the balanced dataset demonstrated better sensitivity and positive predictive value (PPV) values for Class 1, showing a lower number of missed diagnosed gastric cancer patients and fewer erroneous predictions.