<p>The proliferation of memes on social media, blending images with text to convey complex messages, presents unique challenges in content moderation, particularly when identifying harmful content. This paper introduces DecodEM-X, an innovative framework designed to enhance the detection of harmful memes through advanced multimodal analysis. Utilizing the Facebook Hateful Memes dataset, DecodEM-X integrates cutting-edge techniques such as RoBERTa and ResNet50 for robust text and image processing, coupled with a novel cross-attention mechanism that optimizes the synergy between textual and visual modalities. This approach considerably improves the clarification of nuanced, context-dependent interactions within memes. Furthermore, the framework employs SHAP for explainability, ensuring clearness and answerability in automated decision-making processes. Our extensive evaluations demonstrate that DecodEM-X not only outperforms existing baseline models with an accuracy of 89.46% and an AUROC of 90.37% but also provides interpretable insights into the classification decisions. This study highlights the potential of DecodEM-X to set new standards in AI-driven content moderation, promising enhanced reliability and ethical compliance in real-world applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DecodEM-X: advancing multimodal meme moderation with robust AI frameworks

  • Hafiz Muhammad Arslan,
  • Tan Zhenhua

摘要

The proliferation of memes on social media, blending images with text to convey complex messages, presents unique challenges in content moderation, particularly when identifying harmful content. This paper introduces DecodEM-X, an innovative framework designed to enhance the detection of harmful memes through advanced multimodal analysis. Utilizing the Facebook Hateful Memes dataset, DecodEM-X integrates cutting-edge techniques such as RoBERTa and ResNet50 for robust text and image processing, coupled with a novel cross-attention mechanism that optimizes the synergy between textual and visual modalities. This approach considerably improves the clarification of nuanced, context-dependent interactions within memes. Furthermore, the framework employs SHAP for explainability, ensuring clearness and answerability in automated decision-making processes. Our extensive evaluations demonstrate that DecodEM-X not only outperforms existing baseline models with an accuracy of 89.46% and an AUROC of 90.37% but also provides interpretable insights into the classification decisions. This study highlights the potential of DecodEM-X to set new standards in AI-driven content moderation, promising enhanced reliability and ethical compliance in real-world applications.