Multi-frequency Fine-Grained Matching for Audio-Visual Segmentation
摘要
Audio-visual segmentation is a recently proposed task, whose main goal is to locate the target of the sound in the image at the pixel level. In practical scenarios, multiple types of audio can coexist, with different objects emitting sounds at different frequencies. However, existing methods only use single-frequency audio information when fusing audio and visual modalities. Moreover, the process of combining images and audio can be quite rough. Therefore, we propose a multi-frequency fine-grained matching method for multiple sound sources scenario. Firstly, we use short-time Fourier transform (STFT) to extract different frequency spectrograms and input them into the audio encoder to extract multi-frequency audio features. Secondly, multi-frequency audio information serves as a prompt in the pixel decoder stage to guide model segmentation. To obtain high-quality prompts, we use an attention method in the Audio-Visual Matching Module (AVMM) to match visual and audio information. The experiments show that our method has a significant improvement over the baseline and achieves state-of-the-art results on the MS3 benchmark (64.1 mIoU on MS3).