SPADE-Based Dynamic-To-Static Image Translation for Visual Localization of Autonomous Systems
摘要
Visual place recognition aims to match current observation against a set of reference landmarks, so as to determine the location of robots. However, the presence of moving objects in surrounding environments, such as vehicles and pedestrians, hampers the performance of visual place recognition. This paper proposes a multi-modal image translation model to translate dynamic images into static ones, thereby reducing the negative effects of moving objects on visual localization. The proposed method leverages semantic information to guide the recovery of static images from the original dynamic scenes. Specifically, two separate encoder-decoder architectures are built to generate static semantic segmentation maps and static images from the input dynamic images, respectively. Particularly, we build connections from the semantic branch to the image branch at different stages of both the encoder and decoder to fuse semantic features and image features through spatially-adaptive normalization (SPADE) in order to enhance details of the recovered static images and remove dynamic objects. The fused features are utilized by the decoder through skip connections to reconstruct static images. Extensive experimental results demonstrate the superiority of the proposed model in recovering realistically static images.