Hate Speech Classification in Text-Embedded Images: Integrating Ontology, Contextual Semantics, and Vision-Language Representations
摘要
The growing influence of text-embedded images in online communication demands effective strategies for identifying hate speech. The use of hate speech in different contexts makes it necessary to study it in a particular context. Simultaneously, identifying hate speech targets is a crucial research domain as it can offer insights into propagation, impacts, and potential interventions against hate speech. In this article, we address the problem of hate speech detection and target identification in text-embedded images by presenting a comprehensive approach that combines textual and visual cues to accurately detect hate speech and targets within the context of the Russia-Ukraine Crisis. Leveraging a dataset of 4,723 text-embedded images centered around this crisis, we integrate features from the knowledge graph, ontological insights to indicate the presence of hate speech presence, TF-IDF, Named Entity Recognition (NER), and robust vision-language representations. We also provide the rationale behind using different features in our implementation. Our method surpasses existing baselines and methodologies, suggesting the importance of each feature we employ in decision-making.