Comparative Analysis of Preprocessing Techniques for KNN Classification on the Diabetes Dataset
摘要
Preprocessing medical datasets adequately is crucial for improving the accuracy and dependability of machine learning models, especially for healthcare applications. This study uses the popular Diabetes dataset to compare several preprocessing techniques for K-Nearest Neighbors (KNN) classification. Three preprocessing methods—Min-Max Scaling, Standardization, and Robust Scaling—were selected. These techniques are vital for reducing the difficulties of dealing with inconsistent, incomplete, and noisy medical data, which affects the accuracy of classification models' predictions and diagnoses. Each preprocessing approach is thoroughly examined in the research, with special emphasis on the normalizing capabilities of Min-Max Scaling, the resilience of Robust Scaling to outliers, and the standardization accomplished by Standardization. This study uses the KNN algorithm to forecast the occurrence of diabetes mellitus, with accuracy being the primary indicator of performance. Findings from the comparative analysis shed light on the subtle effects of each preprocessing method on the KNN classifier's prediction accuracy, offering helpful information for improving the classification model for diabetes diagnosis. The work aims to further the development of trustworthy and understandable models for healthcare applications by highlighting the need of careful preprocessing of medical datasets.