Listen and Segment: A GNN-Based Network with Attention Mechanism
摘要
Apart from the widely researched detection and segmentation domains, we reconnoitered the embodiment to audio-visual segmentation (AVS) which consists of visual object localization with connected audio patches for each frame. Recently, pixel level image categorization had achieved saturation levels with satisfactory results. But their corresponding audio segments were least explored dropping that valuable support information which makes the model less prone to multiple real-time scenarios. Hence, in this paper, a graphical neural network with attention-based audio-visual segmentation (GWA-AVS) model is proposed, which is a full attention-based graphical neural network (GNN) designed to segment object masks from the audio of the visual frames with single or multiple sound sources. Distinct approaches were investigated to fine-tune the proposed GWA-AVS frameworks to get superior results. Attention block is added to improve the learning of meaningful features of visual frames, thereby improving performance of the proposed framework. GWA-AVS is benchmarked on AVSBench dataset achieving better results than the existing SOTA models with less computational cost, i.e., by using less number of model parameters than existing AVS model. Qualitative results of GWA-AVS evidently show that the generated masks are clear with crisp shape of sounding objects depicting the robustness of our model in real-time scenarios.