Time-Series Clustering of eGFR Data to Enhance Kidney Function Prediction Efficiency
摘要
Deterioration of kidney function is commonly gauged by the decrease in estimated Glomerular Filtration Rate (eGFR), a critical clinical parameter. Numerous machine learning studies have predominantly utilized eGFR values directly. However, this study adopts a different approach, transforming eGFR through time-series clustering into unsorted cluster numbers before incorporating them into the predictive model. Utilizing data from 3,024 participants in the Korean Genomics and Epidemiology Study, spanning 14 years, we propose a comprehensive methodology encompassing data categorization, clustering, and binary marking, followed by the application of Support Vector Machine (SVM) and Ridge Regression for classification. Employing Support Vector Machine (SVM), we achieved promising results with a multi-class prediction accuracy of 0.798, a recall of 0.798, a precision of 0.795, and an F1 score of 0.791. Furthermore, the application of Ridge Regression yielded comparable outcomes, with an accuracy of 0.788, recall of 0.788, precision of 0.755, and an F1 score of 0.740. These outcomes validate that it is feasible to maintain the performance of classification and prediction models without the direct integration of eGFR.