Comprehensive Analysis of Clustering Techniques on Microblog Tweets
摘要
In the digital era, social media has become integral to daily life, generating large amounts of data that has valuable insights on diverse topics. Dealing with these extensive unlabeled datasets has increased the necessity of effective unsupervised data mining techniques. Clustering, a well-known unsupervised technique with wide applications in social media analytics, presents both advantages and disadvantages when applied to real-world and noisy data. This paper conducts a comparative analysis of three widely used clustering techniques: K-Means, HDBSCAN, and Classix. These algorithms are experimented with a novel sensitive dataset, comprising 500 tweets collected for each of five different sensitive tweet hashtags. Hence, the algorithms are evaluated over 2500 tweets to generate five clusters. The performance of these clusters is evaluated using metrics such as Adjusted Rank Index, Adjusted Mutual Information, Homogeneity, Completeness Score, and V measure. Among the three compared algorithms, K-means emerges as the top performer in performance evaluation.