Density-based clustering is a technique that can build clusters with arbitrary shapes while filtering noise, making it useful in diverse applications. The increasing volume of data generated by modern systems has driven the development of density-based clustering algorithms capable of handling large datasets. These algorithms employ methods that reduce object comparisons, reducing runtime without compromising clustering quality. Achieving a good balance between quality and runtime is crucial, as it determines the algorithm’s applicability to real-world problems where both accuracy and runtime are critical. However, to our knowledge, there is no comparison of recent clustering algorithms under the same framework. This work presents an empirical comparison of five density-based clustering algorithms. The comparison framework ensures a fair evaluation by implementing all algorithms with the same data structures and methods. Public standard datasets are used to assess clustering quality and runtime performance. Our results report the algorithm that achieves the best runtime, the highest clustering quality, and a good trade-off between these two factors, providing valuable insights for selecting the most suitable method for large-scale data processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical Comparison of Density-Based Clustering Algorithms for Large Datasets

  • Adrián J. Ramírez-Díaz,
  • José Fco. Martínez-Trinidad,
  • J. Ariel Carrasco-Ochoa

摘要

Density-based clustering is a technique that can build clusters with arbitrary shapes while filtering noise, making it useful in diverse applications. The increasing volume of data generated by modern systems has driven the development of density-based clustering algorithms capable of handling large datasets. These algorithms employ methods that reduce object comparisons, reducing runtime without compromising clustering quality. Achieving a good balance between quality and runtime is crucial, as it determines the algorithm’s applicability to real-world problems where both accuracy and runtime are critical. However, to our knowledge, there is no comparison of recent clustering algorithms under the same framework. This work presents an empirical comparison of five density-based clustering algorithms. The comparison framework ensures a fair evaluation by implementing all algorithms with the same data structures and methods. Public standard datasets are used to assess clustering quality and runtime performance. Our results report the algorithm that achieves the best runtime, the highest clustering quality, and a good trade-off between these two factors, providing valuable insights for selecting the most suitable method for large-scale data processing.