<p>Data clustering is one of the widely explored research problems in the area of machine learning. Data points that are similar to each other are grouped in a cluster using clustering algorithms, whereas data points that are dissimilar are grouped into separate clusters. However, the efficiency of cluster quality compromises with respect to the growing volume of data. In order to address this challenge, researchers are currently moving toward finding a scalable clustering solution. This paper proposes a scalable clustering approach for complex-shaped big data (SCACSBD). It makes use of the initial dataset’s partitioning, applies Density-Based Spatial Clustering of Applications with Noise (DBSCAN) to each partition to obtain partial clustering solutions, and then merges the obtained solutions using Apache spark and Map-Reduce architectures to derive a global clustering solution as the output. The working efficiency of the model is tested with five standard data sets using various clustering validity indices (CVI). Comparative performance analysis reveals that the proposed method outperforms some of the existing state-of-the-art with respect to the CVI values. Asymptotic analysis also shows a comparable computational efficiency of the SCACSBD method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SCACSBD: a scalable clustering approach for complex-shaped big data

  • Agnijit Basu,
  • Binu Jose Ambrose,
  • Raju Hazari,
  • Pranesh Das

摘要

Data clustering is one of the widely explored research problems in the area of machine learning. Data points that are similar to each other are grouped in a cluster using clustering algorithms, whereas data points that are dissimilar are grouped into separate clusters. However, the efficiency of cluster quality compromises with respect to the growing volume of data. In order to address this challenge, researchers are currently moving toward finding a scalable clustering solution. This paper proposes a scalable clustering approach for complex-shaped big data (SCACSBD). It makes use of the initial dataset’s partitioning, applies Density-Based Spatial Clustering of Applications with Noise (DBSCAN) to each partition to obtain partial clustering solutions, and then merges the obtained solutions using Apache spark and Map-Reduce architectures to derive a global clustering solution as the output. The working efficiency of the model is tested with five standard data sets using various clustering validity indices (CVI). Comparative performance analysis reveals that the proposed method outperforms some of the existing state-of-the-art with respect to the CVI values. Asymptotic analysis also shows a comparable computational efficiency of the SCACSBD method.