Efficient clustering approach based on Gower distance for high-dimensional medical datasets
摘要
This paper proposes a novel clustering approach based on Gower distance and a new centroid cluster selection technique. Traditional clustering methods often struggle to improve the performance of standard classifiers when dealing with high-dimensional and heterogeneous data, leading to suboptimal outcomes. Moreover, methods like K-means are highly dependent on the initial selection of centers, which must be predefined, and they fail to effectively identify complex manifold clusters. The introduced approach overcomes these problems by segmenting records into blocks having 2–10 samples to understand their distribution and, further, choose cluster centers in an effective manner. Finally, Gower distance is used to cluster the data based on record similarity. Validation of the proposed clustering method was done using the two medical datasets, the Parkinson’s Disease (PD) and Wisconsin Diagnostic Breast Cancer (WDBC), and Bonn EEG datasets. Results obtained show drastic improvement in machine learning classifiers’ performance, achieving an accuracy of almost 100% for both the PD and WDBC datasets, and 98% for the Bonn dataset. This demonstrates the robustness and effectiveness of the proposed approach compared to traditional methods such as Particle Swarm Optimization (PSO), which requires iterations of optimization. The proposed approach simplifies the clustering process, making it more efficient and less computationally intensive. Implementing this clustering approach in medical organizations can enhance patient care by providing more accurate and reliable data classification. By efficiently handling high-dimensional data, healthcare professionals can make better-informed treatment decisions, improving overall care quality.