错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Audio-Visual Contexts with Refinement for Segmentation

  • Qingwei Geng,
  • Xiaodong Gu

摘要

A more fine-grained video spatial localization task audio visual segmentation(AVS) has recently been proposed, which aims to generate the masks of the sounding objects that sound in the given videos. In this paper, we propose a novel network for AVS. Specifically, to capture comprehensive global auditory semantic information and facilitate its interaction with visual frames, our model incorporates a transformer-based audio-visual context encoder, which is designed to generate a pixel-level score map enriched with auditory contextualized information for the decoder. In addition, to address challenges of the vague boundaries of the sounding object in videos, we introduce a refinement module as well as structural similarity(SSIM) loss to enhance accurate boundary predictions. Extensive experiments on the AVSBench dataset show that our proposed method surpasses the baseline AVSBench as well as some advanced methods from other tasks.