错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mmds: multimodal benchmark dataset for suspicious profile detection on twitter social network

  • Monika Choudhary,
  • Spandan Patil,
  • Satyendra Singh Chouhan,
  • Emmanuel S. Pilli

摘要

In the era of widespread social media usage, detecting and mitigating suspicious profiles is essential for maintaining social platform integrity. While various approaches, including human moderation, machine learning, and network analysis, have been employed, many of the existing frameworks are unimodal in approach. When identifying suspicious users, these approaches typically focus on a single aspect, such as user profile attributes or post-engagement metrics. Little work has been done in leveraging multimodal information, such as text, images, and statistical data associated with user profiles and posts, that can be analyzed collectively to classify suspicious users. This gap can be attributed to the lack of comprehensive multimodal labeled datasets of users on social media. This paper introduces “MmDs”, a multimodal dataset of suspicious and non-suspicious user profiles extracted from the Twitter social network. The curated dataset contains user profile details, timeline information, tweets carrying text and images, associated metadata, and user labels (suspicious/non-suspicious). MmD provides a detailed representation of user behavior and activities. Rigorous validation procedures leveraging statistical tests and clustering metrics have been performed to ensure the quality and relevance of MmDs. Moreover, to demonstrate the practical utility of MmDs, Machine Learning and Deep Learning experiments were conducted on MmDs. The multimodal DL model achieved a high accuracy of 92.24% in identifying suspicious profiles. The baseline experiments show the effectiveness and usefulness of the proposed dataset. Overall, we present a detailed description of MmDs curation and conduct experimental analysis to show the viability of the presented dataset. Our work contributes a valuable resource to the research community engaged in identifying suspicious profiles on social networks.