<p>The rapid growth of data-driven applications has underlined the need for strong methods to analyze and cluster streaming data. Data stream clustering is envisioned to uncover interesting knowledge concealed within data streams, which are typically fast, structure- and pattern-evolving. However, most of the current methods suffer from serious challenges regarding the inability to detect arbitrarily shaped clusters, handling outliers, adaptation to concept drift, and reduction of dependency on predefined parameters. To address these challenges, we propose a novel fractal-dimension-based stream clustering algorithm that detects concept drift and recurrence. Fractal dimension is a versatile mathematical tool that enhances efficiency in clustering the intricate dynamics of evolving streams. Our methodology incorporates a grid-density-based clustering model augmented by dual concept drift detection mechanisms—distance-based metrics and fractal dimensions—to improve sensitivity to subtle distributional shifts. By dynamically updating clusters through model reuse or creation, our algorithm ensures adaptability to real-time changes in data distributions. The proposed algorithm was comprehensively evaluated using the KDDCup-99, Cover Type, and Electricity and Eye State datasets, under diverse scenarios, including concept drifts, evolving data distributions, varying cluster sizes, and outlier conditions. Empirical results demonstrated the algorithm’s superiority over baseline approaches such as StreamKM + + , DenStream, CluStream, and ClusTree, achieving perfect performance metrics. These findings emphasize the effectiveness of our algorithm in addressing real-world streaming data challenges, combining high sensitivity to concept drift with computational efficiency, adaptability, and robust clustering capabilities.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data stream clustering with concept drift using fractal dimension

  • Zahra Rezaei,
  • Hedieh Sajedi,
  • Morteza Hashemi

摘要

The rapid growth of data-driven applications has underlined the need for strong methods to analyze and cluster streaming data. Data stream clustering is envisioned to uncover interesting knowledge concealed within data streams, which are typically fast, structure- and pattern-evolving. However, most of the current methods suffer from serious challenges regarding the inability to detect arbitrarily shaped clusters, handling outliers, adaptation to concept drift, and reduction of dependency on predefined parameters. To address these challenges, we propose a novel fractal-dimension-based stream clustering algorithm that detects concept drift and recurrence. Fractal dimension is a versatile mathematical tool that enhances efficiency in clustering the intricate dynamics of evolving streams. Our methodology incorporates a grid-density-based clustering model augmented by dual concept drift detection mechanisms—distance-based metrics and fractal dimensions—to improve sensitivity to subtle distributional shifts. By dynamically updating clusters through model reuse or creation, our algorithm ensures adaptability to real-time changes in data distributions. The proposed algorithm was comprehensively evaluated using the KDDCup-99, Cover Type, and Electricity and Eye State datasets, under diverse scenarios, including concept drifts, evolving data distributions, varying cluster sizes, and outlier conditions. Empirical results demonstrated the algorithm’s superiority over baseline approaches such as StreamKM + + , DenStream, CluStream, and ClusTree, achieving perfect performance metrics. These findings emphasize the effectiveness of our algorithm in addressing real-world streaming data challenges, combining high sensitivity to concept drift with computational efficiency, adaptability, and robust clustering capabilities.