<p>The rapid expansion of user-generated content on online social networks has intensified the need for accurate and efficient abusive language detection systems. This study proposes a binary classification framework that redefines the widely used Hate Speech and Offensive Language Dataset by merging hate speech and offensive content into a single abusive class, while retaining non-abusive as a separate category. The methodology integrates both semantic and subword-level textual representations by leveraging word-level embeddings from BERT and character-level features from RoBERTa. These heterogeneous features are combined using a multi-neural network model that ensures effective fusion of information across different linguistic granularities. To enhance the discriminative capacity of the model, we incorporate a feature enhancement module comprising CNNs for local pattern extraction, Bidirectional Long Short-Term Memory (Bi-LSTM) networks for capturing sequential dependencies, and a self-attention mechanism to model global contextual relationships. The final classification is performed through a sigmoid-activated dense layer trained using binary cross-entropy loss. Experimental evaluations show that the proposed hybrid architecture significantly outperforms models based solely on word-level or character-level features, achieving higher accuracy, precision, recall, and F1-score. The results confirm that the combination of multi-granular features with contextual enhancement offers a robust and scalable solution for detecting abusive content in social media environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-granular hybrid neural architecture for detecting abusive content in online social networks (OSNs) with contextual awareness

  • Uday Shankar Yadavalli,
  • Somya Ranjan Sahoo

摘要

The rapid expansion of user-generated content on online social networks has intensified the need for accurate and efficient abusive language detection systems. This study proposes a binary classification framework that redefines the widely used Hate Speech and Offensive Language Dataset by merging hate speech and offensive content into a single abusive class, while retaining non-abusive as a separate category. The methodology integrates both semantic and subword-level textual representations by leveraging word-level embeddings from BERT and character-level features from RoBERTa. These heterogeneous features are combined using a multi-neural network model that ensures effective fusion of information across different linguistic granularities. To enhance the discriminative capacity of the model, we incorporate a feature enhancement module comprising CNNs for local pattern extraction, Bidirectional Long Short-Term Memory (Bi-LSTM) networks for capturing sequential dependencies, and a self-attention mechanism to model global contextual relationships. The final classification is performed through a sigmoid-activated dense layer trained using binary cross-entropy loss. Experimental evaluations show that the proposed hybrid architecture significantly outperforms models based solely on word-level or character-level features, achieving higher accuracy, precision, recall, and F1-score. The results confirm that the combination of multi-granular features with contextual enhancement offers a robust and scalable solution for detecting abusive content in social media environments.