A Novel Framework Integrating Epidemiological Features and Machine Learning Models for Prostate Cancer Detection on Imbalanced Dataset
摘要
Prostate cancer represents a major health issue worldwide, where traditional biopsy-based diagnostics are often limited by patient discomfort, bias, subjectivity, and sampling inaccuracies. This study strives to mitigate these challenges by utilizing multiple machine learning algorithms to detect prostate cancer early by utilizing diverse epidemiological features to minimize geographic and racial disparities. Notably, the dataset contains a diverse set of epidemiological and screening-related features, while let us have comprehensive evaluation of risk-associated variables in prostate cancer detection. This study introduces the ProstaEnsembleNet framework, which incorporates machine learning models in an ensemble approach. The designed ensemble learning integrates machine learning models, with logistic regression serving as the meta-learner. In addition, gradient boosting, random forest, Gaussian Naive Bayes, extreme gradient boosting, support vector machines, light gradient boosting machines, and k nearest neighbor, alongside deep learning models such as TabNet and multilayer perceptron are utilized for experiments. To address data imbalance, the synthetic minority oversampling technique is applied to the training dataset. Using a prostate cancer dataset containing 29 clinical attributes, the experimental analysis revealed that the proposed ensemble approach achieved superior performance by scoring an F1 score of 91.42%, a recall of 98.02%, and an PR-AUC of 86.32%. The results underscore the effectiveness of class-balancing strategies and ensemble learning approaches in enhancing diagnostic accuracy within imbalanced medical datasets.