Audio-Visual Emotion Recognition (AVER) is essential in various real-world applications. Many methods try to extract and fuse the audio and visual modalities to comprehend better and classify the underlying emotion. Recently, large pre-trained models brought powerful modality-fusion ability in general datasets and significantly outperformed traditional small-scale models. However, they are less effective in complementing some specialized scenarios due to the conflict of meanings between the two modalities. This paper proposes a parameter-efficient fine-tuning method, Content-Aware Efficient Learner (CAEL), to solve this problem with subtle computational consumption. Specifically, we propose an adapter network based on the pre-trained audio and visual transformers for modality fusion. To better fuse the two modalities, propose content-aware attention, in which the audio and visual information align and fuse under the guidance of the speech content. Extensive experiments on the CREMA-D dataset verify the effectiveness and efficiency of our proposed framework.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Content-Aware Efficient Learner for Audio-Visual Emotion Recognition

  • Guanjie Huang,
  • Weilin Lin,
  • Li Liu

摘要

Audio-Visual Emotion Recognition (AVER) is essential in various real-world applications. Many methods try to extract and fuse the audio and visual modalities to comprehend better and classify the underlying emotion. Recently, large pre-trained models brought powerful modality-fusion ability in general datasets and significantly outperformed traditional small-scale models. However, they are less effective in complementing some specialized scenarios due to the conflict of meanings between the two modalities. This paper proposes a parameter-efficient fine-tuning method, Content-Aware Efficient Learner (CAEL), to solve this problem with subtle computational consumption. Specifically, we propose an adapter network based on the pre-trained audio and visual transformers for modality fusion. To better fuse the two modalities, propose content-aware attention, in which the audio and visual information align and fuse under the guidance of the speech content. Extensive experiments on the CREMA-D dataset verify the effectiveness and efficiency of our proposed framework.