Harmful memes contain multimodal content that integrates text and images and conveys malicious information. The task of harmful meme detection determines whether a meme is harmful, while the task of harmful meme identification further categorizes the harmful content into specific subtypes. Despite the progress made in research on harmful memes, the following challenges still exist in research on Chinese harmful memes. (i) Reliance on distinct cultural contexts, colloquial expressions, and internet buzzwords makes it difficult to accurately extract harmful cues. (ii) frequent use of homophones and metaphors obscures the intended meaning; and (iii) inherent semantic ambiguity arising from humor, sarcasm, and self-deprecation further complicates harmfulness assessment. To address the above challenges, we propose a Graph Embedding-Enhanced Textual Inversion framework (GEETI) to solve the Chinese harmful meme detection and identification task. First, we propose a cross-modal emotion and semantic parser module (CESP) to address the difficulties posed by Chinese harmful memes that rely on distinct cultural contexts, colloquial expressions, and internet buzzwords. Then, to resolve the ambiguity introduced by homophones and metaphors, we employ a textual inversion branch powered by a graph attention network to fuse visual and textual features and generate pseudo-text representations. Finally, to mitigate the intrinsic semantic ambiguity caused by humor, sarcasm, and self-deprecation, we incorporate a feature interaction matrix to capture fine-grained cross-modal associations, thus enhancing alignment fusion and improving classification accuracy. Experimental results on the publicly available TOXICN MM dataset demonstrate that GEETI achieves state-of-the-art performance. Our code for this work is available at https://github.com/fus0618/GEETI . Disclaimer: This paper contains discriminatory content that may not be comforting to some readers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GEETI: Graph Embedding-Enhanced Textual Inversion for Chinese Harmful Meme Detection and Identification

  • Shichao Fu,
  • Tongguan Wang,
  • Ying Sha

摘要

Harmful memes contain multimodal content that integrates text and images and conveys malicious information. The task of harmful meme detection determines whether a meme is harmful, while the task of harmful meme identification further categorizes the harmful content into specific subtypes. Despite the progress made in research on harmful memes, the following challenges still exist in research on Chinese harmful memes. (i) Reliance on distinct cultural contexts, colloquial expressions, and internet buzzwords makes it difficult to accurately extract harmful cues. (ii) frequent use of homophones and metaphors obscures the intended meaning; and (iii) inherent semantic ambiguity arising from humor, sarcasm, and self-deprecation further complicates harmfulness assessment. To address the above challenges, we propose a Graph Embedding-Enhanced Textual Inversion framework (GEETI) to solve the Chinese harmful meme detection and identification task. First, we propose a cross-modal emotion and semantic parser module (CESP) to address the difficulties posed by Chinese harmful memes that rely on distinct cultural contexts, colloquial expressions, and internet buzzwords. Then, to resolve the ambiguity introduced by homophones and metaphors, we employ a textual inversion branch powered by a graph attention network to fuse visual and textual features and generate pseudo-text representations. Finally, to mitigate the intrinsic semantic ambiguity caused by humor, sarcasm, and self-deprecation, we incorporate a feature interaction matrix to capture fine-grained cross-modal associations, thus enhancing alignment fusion and improving classification accuracy. Experimental results on the publicly available TOXICN MM dataset demonstrate that GEETI achieves state-of-the-art performance. Our code for this work is available at https://github.com/fus0618/GEETI . Disclaimer: This paper contains discriminatory content that may not be comforting to some readers.