Spatiotemporal information cooperative interaction network for video salient object detection
摘要
Video salient object detection is all about identifying the most salient object from a video sequence and segmenting the exact region of that object. Most video salient object detection methods have low performance and low efficiency, and there is much room for improvement. We found that some existing methods consider the advantages of spatial and temporal modalities, but do not fully explore the complementarity between different modalities, which may lead to poor performance of the model when dealing with complex scenes. To cope with the above problems, this paper proposes a novel end-to-end spatiotemporal information cooperative interaction network (SICINet) for salient object detection in video. The network consists of three key modules: cross-modal feature supplementation (CFS), cross-guidance enhancement (CGE), and refinement-adaptive fusion (RAF). Specifically, we propose the CFS module to enable spatiotemporal features to complement each other to facilitate discriminative feature learning, and to learn cross-modal feedback features to ensure the comprehensiveness of salient information in the subsequent fusion stage. In addition, considering that spatiotemporal information can be mutually constrained, we designed the CIE module to use the rough prediction maps of spatiotemporal modalities to bootstrap each other’s multilevel features to filter the unimodal redundant information. Finally, we introduce the RAF module to refine the input features using spatial attention and achieve adaptive weighting fusion by learning the channel weights. Experimental evaluations on four publicly available datasets show that our proposed method is robust under various challenging scenarios (e.g., multiple objects, dynamic foregrounds) and performs more favorably than current state-of-the-art methods.