Sarcasm Detection: A Multi-Modal Approach Integrating Text, Audio, and Visual Cues
摘要
Sarcasm detection has been an area of increasing interest in Natural Language Processing (NLP) due to the unique linguistic properties of sarcasm, which make it a challenging task for machines. Traditionally, sarcasm detection has focused on text-only datasets. However, relying solely on text can limit the understanding of the nuances of sarcasm, which often involves vocal and visual cues. This paper proposes a multi-modal approach integrating text, audio, and visual data to enhance sarcasm detection accuracy. By using a combination of deep learning techniques applied to multi-modal data, this work demonstrates significant improvements in sarcasm detection performance. We introduce new datasets combining text from social media, vocal data from video platforms, and visual data to improve sarcasm interpretation, achieving an accuracy rate of over 91.4%.