Target-Oriented Dynamic Denosing Curriculum Learning for Multimodel Stance Detection
摘要
Multimodal Stance Detection aims to classify public opinions on specific targets in social media, incorporating both text and image data. However, prior studies have overemphasized the significance of images, neglecting the presence of irrelevant images in the dataset. Moreover, previous research has shown that employing the Chain of Thought approach with large language models can introduce noise from the generated text as well as noise from the image modality into the text input modality. These noises can degrade the performance of multimodal models. Additionally, both image and text modalities exhibit complex data patterns, resulting in significant disparities in training difficulty across the dataset. To address these issues, We proposed Target-Oriented Dynamic Denoising Curriculum Learning(TODDCL), which effectively measures different types of noise and automatically tackles the noise in both image and text modalities based on their relevance to the target, in an escalating complexity order. Experimental results on five benchmark datasets demonstrate that our proposed TODDCL method achieves state-of-the-art performance in multimodal stance detection.