<p>Rating prediction is an important task in review mining. Aspect category sentiment analysis (ACSA) and overall rating prediction (ORP) are closely related tasks, yet existing joint models are typically limited to text-only inputs, while existing multimodal review models usually address only a single task. To bridge this gap, we propose MMRP, an end-to-end multimodal multitask framework for jointly modeling ACSA and ORP from image-text reviews. The model integrates textual and visual information through a cross-modal Transformer and further introduces modality-aware learnable positional encoding and a gated multi-head cross-modal attention mechanism to better handle the structure of multimodal reviews, where text forms token sequences and images appear as variable-sized sets. We evaluate MMRP on the ZOL mobile review dataset and an additional cross-domain multimodal benchmark, ViMACSA. Extensive experiments, including comparative evaluation, ablation analysis, sensitivity analysis, and case studies, show that MMRP consistently achieves strong performance and exhibits good robustness, cross-domain applicability, and practical deployment potential. These results demonstrate the effectiveness of jointly leveraging multimodal information for aspect-level and overall rating prediction in user reviews.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint aspect-based sentiment and overall rating prediction via a cross-modal transformer in user reviews

  • Xiaoning Wang,
  • Rui Li,
  • Yuxuan Guo,
  • Xiaoling Lu,
  • Jiawei Ren,
  • Libo Sun

摘要

Rating prediction is an important task in review mining. Aspect category sentiment analysis (ACSA) and overall rating prediction (ORP) are closely related tasks, yet existing joint models are typically limited to text-only inputs, while existing multimodal review models usually address only a single task. To bridge this gap, we propose MMRP, an end-to-end multimodal multitask framework for jointly modeling ACSA and ORP from image-text reviews. The model integrates textual and visual information through a cross-modal Transformer and further introduces modality-aware learnable positional encoding and a gated multi-head cross-modal attention mechanism to better handle the structure of multimodal reviews, where text forms token sequences and images appear as variable-sized sets. We evaluate MMRP on the ZOL mobile review dataset and an additional cross-domain multimodal benchmark, ViMACSA. Extensive experiments, including comparative evaluation, ablation analysis, sensitivity analysis, and case studies, show that MMRP consistently achieves strong performance and exhibits good robustness, cross-domain applicability, and practical deployment potential. These results demonstrate the effectiveness of jointly leveraging multimodal information for aspect-level and overall rating prediction in user reviews.