Person Re-Identification with CNN ResNet Spatial–Temporal Stream
摘要
Identifying individuals from many cameras that correspond to a given query is the goal of video-based re-identification of people (reID). Video-based person re-identification makes considerable use of the temporal attention mechanism. This challenge is significantly more difficult than image-based person reID because of the spatial and temporal distractions present in person videos, which include backdrop noise and partial closures over frames, respectively. We see that spatial distracting factors always appear in a particular place, while temporal distracting factors show different patterns, such as partial closures in the initial few frames. These patterns provide helpful cues about which frames or temporal attentions to pay attention to first. Therefore, this study works with ResNet101 as the backbone network while using spatial–temporal stream to provide an improved person re-identification method using YOLOv5 which effectively extracts features from the input source and provides a lightweight model.