<p>Nowadays, online social networks are continuously evolving and are utilized across a multitude of applications. Nevertheless, these networks are not immune to challenges, as they are susceptible to various forms of social spam, which refers to unsolicited or unwanted content or messages on social networking platforms. To mitigate the effects of such malicious activities, we introduce in this paper a new system that effectively detects social spam. Our approach utilizes a novel features representation method based on a proposed Transformers Features Curve (TFC), which allows us to generate contextualized embeddings and extract highly informative features to enhance the representation of content features, which is a limitation in existing work. Features are subsequently extracted and classified from the generated 2D curve representation using multiple architectural configurations of the 2D Convolutional Neural Network (CNN), which gives rise to TFC-CNN, the name of our architecture. The experiments have shown that this new classification technique has outperformed Machine Learning algorithms and existing works using two different textual spam datasets. In Arabic, the model achieved an impressive accuracy of 98.68%, with a precision of 99.99%, recall of 98.32%, and an F1 score of 99.15%. Additionally, for English, the model demonstrated an accuracy of 93.67%, a precision of 94.26%, a recall of 94.07%, and an F1 score of 94.17%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient arabic and english social spam detection using a transformer and 2D convolutional neural network-based deep learning filter

  • Marouane Kihal,
  • Lamia Hamza

摘要

Nowadays, online social networks are continuously evolving and are utilized across a multitude of applications. Nevertheless, these networks are not immune to challenges, as they are susceptible to various forms of social spam, which refers to unsolicited or unwanted content or messages on social networking platforms. To mitigate the effects of such malicious activities, we introduce in this paper a new system that effectively detects social spam. Our approach utilizes a novel features representation method based on a proposed Transformers Features Curve (TFC), which allows us to generate contextualized embeddings and extract highly informative features to enhance the representation of content features, which is a limitation in existing work. Features are subsequently extracted and classified from the generated 2D curve representation using multiple architectural configurations of the 2D Convolutional Neural Network (CNN), which gives rise to TFC-CNN, the name of our architecture. The experiments have shown that this new classification technique has outperformed Machine Learning algorithms and existing works using two different textual spam datasets. In Arabic, the model achieved an impressive accuracy of 98.68%, with a precision of 99.99%, recall of 98.32%, and an F1 score of 99.15%. Additionally, for English, the model demonstrated an accuracy of 93.67%, a precision of 94.26%, a recall of 94.07%, and an F1 score of 94.17%.