A prompt-based dual-layer cross-modal distillation learning method for aspect-based sentiment analysis
摘要
Multimodal aspect-based sentiment analysis (MABSA) identifies the sentiment polarity of specific aspects in text with the aid of accompanying images. Current methods often exhibit limited robustness to visual noise and scene adaptability, particularly under visually noisy conditions or image absence, leading to suboptimal performance. To address this, we propose a Prompt-based Dual-layer Cross-modal Distillation (PDCD) learning strategy. PDCD leverages aspect term prompts to focus the model on relevant aspect semantics and employs a dual-layer cross-modal distillation mechanism integrating textual and visual features. This approach progressively extracts valuable visual cues while suppressing noise, enabling accurate aspect sentiment prediction even with missing images or mismatched image-text pairs. The proposed PDCD offers two key innovations: (1) its dual-layer cross-modal distillation framework demonstrates strong adaptability to diverse scenarios and noise types. (2) Its prompt-guided dual-aspect representation effectively aggregates aspect semantics within noisy visual conditions. Experiments on Twitter2015, Twitter2017, and MASAD benchmarks show PDCD outperforms most existing state-of-the-art (SOTA) methods by 75.69 F1-score on average. Additional visual noise tests confirm superior robustness across diverse noise scenarios, achieving the highest F1-scores. In-depth analyses further validate the strategy’s effectiveness in noisy MABSA scenarios.