错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Twitter spam drift detection by semi supervised learning approach using YATSI algorithm

  • P. Sivakumar,
  • M. Balasubramani,
  • R. Sowndharya,
  • B. S. Deepa Priya,
  • W. Deva Priya,
  • Maganti Syamala

摘要

Twitter has improved in such a way people acquire knowledge or information by making them share their thoughts and opinions on everyday tweets. However, spammers have discovered Twitter to be desirable for spreading spam as a result of its enormous popularity. Twitter spam, in contrast to other types of spam, has recently become a big concern. The enormous number of users and volume of content or information published on Twitter contribute considerably to the rise of spam. To protect users, Twitter and the research team have developed several spam detection systems that employ various machine-learning techniques. According to a new study, existing machine learning-based detection algorithms are unable to detect spam correctly since the features of spam tweets vary over time. The issue is referred to as “Twitter Spam Drift.” In this paper, a semi-supervised learning approach (SSLA) using the YATSI algorithm has been suggested. YATSI is categorized into two steps. An initial prediction model is the first phase. The genuine predictions for unlabeled cases are identified in the second phase by using ML algorithms. To deal with the drift, the study utilizes a live Twitter stream of data acquired using Twitter API. This proposed method uses pre-processed labelled data to learn the structure of unlabeled data that is live-downloaded to distinguish between genuine and fake users. Experiments were conducted on live twitter data using KNN, SVM and NB machine learning classifiers. Among those classifiers SVM is showing the better results, in-terms of accuracy.