Stacking ensemble model for predicting chronic kidney disease in the Uddanam region of India with unknown etiology
摘要
Chronic Kidney Disease (CKD) is a severe public health issue across the globe. It is hard to detect early because there are no apparent symptoms. Many patients do not realize they have the disease. Improved models are crucial for altering early diagnosis and effective treatment. No single cause has been found to predict the disease. This study uses a new dataset from Uddanam villages in Srikakulam district, AP, India, where CKD is common. Data preprocessing was done carefully to maintain data integrity. Outliers were detected and removed to improve dataset quality. Exploratory Data Analysis (EDA) techniques were used to analyze the processed data. It included statistical analysis, feature analysis, target attributes, and outlier detection analysis. In this study, we collected patients’ data from various clinical centers and hospitals in the Uddanam area of Andhra Pradesh, India. The study uses a Stacking Ensemble method to improve predictions of chronic kidney disease in the Uddanam region. It combines Naïve Bayes, K-NN, and SVM with Sigmoid and Linear kernels. Logistic Regression is used as the meta-learner. We use principal component analysis (PCA) for dimensional reduction to better combine features of CKD analysis and prediction with ML models. The PCA with the Stacking Ensemble approach verified good performance compared to the state-of-the-art models, achieving 98.9% on CA, 98.9% on F1, and 0.999 on AUC. It has higher predictive accuracy and better generalizability across the dataset. The “Uddanam Chronic Kidney Disease (UCKD)” study shows that advanced ML techniques can create effective models for CKD. The results highlight the need to collect and prepare local data. The study helps improve how well machine learning works in healthcare.