Integrating Audio-Visual Contexts with Refinement for Segmentation
摘要
A more fine-grained video spatial localization task audio visual segmentation(AVS) has recently been proposed, which aims to generate the masks of the sounding objects that sound in the given videos. In this paper, we propose a novel network for AVS. Specifically, to capture comprehensive global auditory semantic information and facilitate its interaction with visual frames, our model incorporates a transformer-based audio-visual context encoder, which is designed to generate a pixel-level score map enriched with auditory contextualized information for the decoder. In addition, to address challenges of the vague boundaries of the sounding object in videos, we introduce a refinement module as well as structural similarity(SSIM) loss to enhance accurate boundary predictions. Extensive experiments on the AVSBench dataset show that our proposed method surpasses the baseline AVSBench as well as some advanced methods from other tasks.