Modal and scenario noise reduction for multimodal sarcasm detection
摘要
Multimodal sarcasm detection integrates textual, acoustic, and visual cues to identify sarcastic intent. However, a counter-intuitive pattern has emerged: state-of-the-art models often achieve better performance by excluding non-textual context, suggesting that naive integration introduces more interference than useful signal. We identify two factors underlying this degradation: modal noise from low-information content in acoustic and visual streams, and scenario noise from inherent ambiguity where identical features indicate opposite intents across different contexts. We propose MSNR-MSD, which addresses these challenges through complementary mechanisms. For modal noise, Dual-layer Modal Noise Reduction (DMNR) decomposes context refinement into redundancy filtering and target-driven fusion, enabling selective preservation of prosodic and emotional cues while removing interference. Empirical analysis reveals that redundancy filtering benefits acoustic but harms visual features, leading to modality-specific processing. For scenario noise, a lightweight calibration module leverages cross-sample prototypes to disambiguate high-uncertainty predictions. Experiments on MUStARD and MUStARD++ benchmarks demonstrate state-of-the-art performance, with ablation studies confirming that both mechanisms contribute consistent gains.