<p>The exponential growth of biomedical data arising from medical images, genomics, bioinformatics, and electronic health records has intensified the urgent need for more scalable, interpretable, and computationally efficient unsupervised learning approaches. Among these, the K-means clustering technique remains one of the most widely adopted clustering methods because of its conceptual, design and implementation simplicity, adaptability, and effectiveness across diverse biomedical data modalities. This paper presents a systematic and critical survey of K-means clustering methodologies applied to biomedical data science, with specific emphasis on scalability, robustness, and integration into modern big data ecosystems. Similarly, this study categorises and analyses classical, enhanced, and hybrid K-means variants, including initialisation strategies, distance metrics, constraint-based formulations, and distributed implementations for high-dimensional, large-scale biomedical datasets. Moreover, the survey also examines recent trends in augmented K-means workflows with large language models (LLMs), highlighting their emerging roles in data preprocessing, feature engineering, semantic annotation, and a human-in-the-loop analytical pipeline rather than in direct numerical optimisation. Lastly, open challenges related to data heterogeneity, interpretability, and scalability, as well as benchmarking practices, are discussed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From K-means to LLM-augmented pipelines: a critical survey and taxonomy for biomedical big data

  • Absalom E. Ezugwu,
  • Joy Ezugwu

摘要

The exponential growth of biomedical data arising from medical images, genomics, bioinformatics, and electronic health records has intensified the urgent need for more scalable, interpretable, and computationally efficient unsupervised learning approaches. Among these, the K-means clustering technique remains one of the most widely adopted clustering methods because of its conceptual, design and implementation simplicity, adaptability, and effectiveness across diverse biomedical data modalities. This paper presents a systematic and critical survey of K-means clustering methodologies applied to biomedical data science, with specific emphasis on scalability, robustness, and integration into modern big data ecosystems. Similarly, this study categorises and analyses classical, enhanced, and hybrid K-means variants, including initialisation strategies, distance metrics, constraint-based formulations, and distributed implementations for high-dimensional, large-scale biomedical datasets. Moreover, the survey also examines recent trends in augmented K-means workflows with large language models (LLMs), highlighting their emerging roles in data preprocessing, feature engineering, semantic annotation, and a human-in-the-loop analytical pipeline rather than in direct numerical optimisation. Lastly, open challenges related to data heterogeneity, interpretability, and scalability, as well as benchmarking practices, are discussed.