Joint aspect-based sentiment and overall rating prediction via a cross-modal transformer in user reviews
摘要
Rating prediction is an important task in review mining. Aspect category sentiment analysis (ACSA) and overall rating prediction (ORP) are closely related tasks, yet existing joint models are typically limited to text-only inputs, while existing multimodal review models usually address only a single task. To bridge this gap, we propose MMRP, an end-to-end multimodal multitask framework for jointly modeling ACSA and ORP from image-text reviews. The model integrates textual and visual information through a cross-modal Transformer and further introduces modality-aware learnable positional encoding and a gated multi-head cross-modal attention mechanism to better handle the structure of multimodal reviews, where text forms token sequences and images appear as variable-sized sets. We evaluate MMRP on the ZOL mobile review dataset and an additional cross-domain multimodal benchmark, ViMACSA. Extensive experiments, including comparative evaluation, ablation analysis, sensitivity analysis, and case studies, show that MMRP consistently achieves strong performance and exhibits good robustness, cross-domain applicability, and practical deployment potential. These results demonstrate the effectiveness of jointly leveraging multimodal information for aspect-level and overall rating prediction in user reviews.