LAMB: Label-Induced Mixed-Level Blending for Multimodal Multi-label Emotion Detection
摘要
To better understand complex human emotions, there is growing interest in utilizing heterogeneous sensory data to detect multiple co-occurring emotions. However, existing studies have focused on extracting static information from each modality, while overlooking various interactions within and between modalities. Additionally, the label-to-modality and label-to-label dependencies still lack exploration. In this paper, we propose LAbel-induced Mixed-level Blending (LAMB) to address these challenges. Mixed-level blending leverages shallow but manifold self-attention and cross-attention encoders in parallel to model unimodal context dependency and cross-modal interaction simultaneously. This is in contrast to previous works either use one of them or cascade them successively, which ignores the diversity of interaction in multimodal data. LAMB also employs label-induced aggregation to allow different labels to attend to the most relevant blended tokens adaptively using a transformer-based decoder, which facilitates the exploration of label-to-modality dependency. Unlike common low-order strategies in multi-label learning, correlations among multiple labels can be learned by self-attention in label embedding space before being treated as queries. Comprehensive experiments demonstrate the effectiveness of our methods for multimodal multi-label emotion detection.