Content-Aware Efficient Learner for Audio-Visual Emotion Recognition
摘要
Audio-Visual Emotion Recognition (AVER) is essential in various real-world applications. Many methods try to extract and fuse the audio and visual modalities to comprehend better and classify the underlying emotion. Recently, large pre-trained models brought powerful modality-fusion ability in general datasets and significantly outperformed traditional small-scale models. However, they are less effective in complementing some specialized scenarios due to the conflict of meanings between the two modalities. This paper proposes a parameter-efficient fine-tuning method, Content-Aware Efficient Learner (CAEL), to solve this problem with subtle computational consumption. Specifically, we propose an adapter network based on the pre-trained audio and visual transformers for modality fusion. To better fuse the two modalities, propose content-aware attention, in which the audio and visual information align and fuse under the guidance of the speech content. Extensive experiments on the CREMA-D dataset verify the effectiveness and efficiency of our proposed framework.