The article introduces a novel approach for identifying and classifying spammer groups in online review systems. The focus is integrating review content and group behavioural features to enhance spam detection accuracy. The proposed method utilizes the Frequent Pattern (FP) Growth Algorithm to generate candidate groups. Subsequently, it analyzes various spammer indicators, including Rating Features, Date Feature, Rating Variance, Group Size, Group Member Content Similarity, Weighted Rating Average, and Review Length Feature. This article employed a Yelp dataset containing reviews of hotels and restaurants and performed data preprocessing techniques for text analysis. Visualization of rating distribution, word frequency, and the relationship between review length and rating is presented to provide insights into the dataset. The proposed method is compared with existing models, demonstrating superior performance across multiple evaluation metrics, achieving an accuracy of 98.36% on Random Forest. The article highlights the significance of the developed approach in detecting fake reviews, particularly those posted by professional spammers. Future directions include exploring unsupervised methods for spammer group detection to improve robustness and scalability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spammer Groups Detection in Online Reviews: A Novel Approach Using FP-Growth and Behavioral Features

  • Arvind Mewada,
  • Sushil Kumar Maurya

摘要

The article introduces a novel approach for identifying and classifying spammer groups in online review systems. The focus is integrating review content and group behavioural features to enhance spam detection accuracy. The proposed method utilizes the Frequent Pattern (FP) Growth Algorithm to generate candidate groups. Subsequently, it analyzes various spammer indicators, including Rating Features, Date Feature, Rating Variance, Group Size, Group Member Content Similarity, Weighted Rating Average, and Review Length Feature. This article employed a Yelp dataset containing reviews of hotels and restaurants and performed data preprocessing techniques for text analysis. Visualization of rating distribution, word frequency, and the relationship between review length and rating is presented to provide insights into the dataset. The proposed method is compared with existing models, demonstrating superior performance across multiple evaluation metrics, achieving an accuracy of 98.36% on Random Forest. The article highlights the significance of the developed approach in detecting fake reviews, particularly those posted by professional spammers. Future directions include exploring unsupervised methods for spammer group detection to improve robustness and scalability.