错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Systematic Framework for Testing Multimodal Feature Contribution in Recommendation Systems

  • Anunobi Victor Chibueze,
  • Jianfang Wang

摘要

Multimodal recommendation systems are presumed to outperform unimodal ones by leveraging diverse data signals (text, images, etc.). However, we find that a state-of-the-art model achieves over 98% of its performance using text alone, rendering its visual modality practically useless. This exposes a widespread but overlooked problem: “pseudo-multimodality”. To address this, we introduce the Systematic Multimodal Ablation Framework (SMAF), the first reproducible protocol for quantifying individual modality contributions through controlled ablation with statistical significance testing across multiple datasets. The SMAF employs three phases: (i) controlled single-modality isolation via feature masking, (ii) systematic modality combination testing, and (iii) cross-dataset validation with statistical rigor. Applying SMAF to SEA, a state-of-the-art multimodal recommender on three Amazon datasets, revealed that SEA achieved over 98% performance using only text features, with image contributions below 2%. This exposes SEA as “pseudo-multimodal” architecturally multimodal but functionally unimodal, attributed to low-quality image features (91% zero vectors) and inadequate concatenation fusion. By challenging the fundamental assumptions about the inherent benefits of multimodality, our findings underscore the need for a paradigm shift in how future multimodal systems are designed. Specifically, they suggest a more critical evaluation of modality contributions and a potential rethinking of resource allocation towards modality that provides the most significant performance gains.