Online shopping platforms are growing at an alarming rate. Consequently, there has been a rapid increase in fake reviews that threaten the authenticity and integrity of the internet commerce. We propose a novel methodology to address this problem that derives semantically rich word embeddings from review text using BERT. These embeddings can then be used as features for training various ML classifier models such as Support Vector Machine (SVM), Random Forest, Naive Bayes, Bagging Classifier, K-Nearest Neighbors (KNN) and AdaBoost. Cornell University dataset, Chicago review dataset, doctor review dataset and Amazon review dataset were employed in evaluating our approach performance wise. SVM classifiers with BERT-generated embeddings outperformed other classifiers with 88.91% accuracy which was 8.7% higher than the state-of-the-art method and showed significant improvement in accuracy and F1 scores in restaurant and doctor datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient and Automated Fake Review Detection with BERT Embeddings

  • Furqan Yaqub Khan,
  • Tawseef Ayoub Shaikh

摘要

Online shopping platforms are growing at an alarming rate. Consequently, there has been a rapid increase in fake reviews that threaten the authenticity and integrity of the internet commerce. We propose a novel methodology to address this problem that derives semantically rich word embeddings from review text using BERT. These embeddings can then be used as features for training various ML classifier models such as Support Vector Machine (SVM), Random Forest, Naive Bayes, Bagging Classifier, K-Nearest Neighbors (KNN) and AdaBoost. Cornell University dataset, Chicago review dataset, doctor review dataset and Amazon review dataset were employed in evaluating our approach performance wise. SVM classifiers with BERT-generated embeddings outperformed other classifiers with 88.91% accuracy which was 8.7% higher than the state-of-the-art method and showed significant improvement in accuracy and F1 scores in restaurant and doctor datasets.