Multimodal Sarcasm Detection Based on Cross-Modal Fine-Grained Fusion and Reasoning Chain Enhancement
摘要
Multimodal sarcasm detection is of great significance for understanding users’ intentions and emotions. Existing methods suffer from weaknesses in visual modality representation, insufficient granularity in cross-modal fusion, and a lack of semantic reasoning knowledge. This paper proposes a multimodal sarcasm detection model based on cross-modal fine-grained fusion and reasoning chain enhancement, conducting research from three dimensions: modal modeling, knowledge-enhanced reasoning, and semantic fusion. (1) We utilize multimodal large language models to perform deep semantic parsing of visual modality images; (2) based on the long-chain reasoning technology of large language models, we achieve knowledge enhancement under the constraints of semantic and logical consistency; (3) we design an image region-text token-level association mapping to achieve more detailed and accurate cross-modal information fusion. Experimental results on the HFM and SarcNet datasets show that the method proposed in this paper has significant advantages over baseline methods.