Diabetes Risk Prediction: A Comparative Analysis of Feature Selection Techniques for Efficient k-means Clustering
摘要
Diabetes, a prevalent and serious health concern, has seen rising incidence globally. This study aims to develop a predictive model for early diabetes detection using machine learning. By leveraging data from an ETL-populated data warehouse, this research explored various feature selection techniques, including gradient boosting classifier (GBC), principal component analysis (PCA), and permutation importance (PI). The effectiveness of K-means clustering is evaluated using metrics such as Silhouette, Calinski-Harabasz, and Davies-Bouldin scores. Our findings highlight the critical role of feature selection in enhancing model performance and efficiency.