In the field of Person re-identification, text-guided semantic information mining has shown strong potential in recent years. The visual language model CLIP, with its excellent cross-modal comprehension capability, has achieved remarkable results in several tasks. In particular, in person re-identification, a two-stage CLIP-based approach has recently been proposed, which attempts to generate corresponding textual descriptions from the labels of specific pedestrians, which in turn aids the pedestrian classification task. However, after in-depth research, it was found that this approach has limitations in the first stage of feature extraction, i.e., the generated features are not sufficient to fully fit the pedestrian’s characteristics, which leads to the fact that the textual cues generated by utilizing these features in the second stage are not able to accurately represent the information for pedestrian in the vision. To solve this problem, this paper proposes an innovative Multi-stage Channel feature Aggregation method. The method digs deeper into the feature information of different stages through a well-designed feature aggregation strategy. This design allows our method to learn feature representations that are more attuned to pedestrian characteristics, which in turn generates more accurate textual cues. On the MSMT17, Market-1501, DukeMTMC, and Occluded-Duke datasets, we achieve accuracy improvements of 2.4%, 1.0%, 1.5%, and 3.6%, respectively. These results fully validate validity and applicability of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Person Re-identification with Multi-stage Channel Feature Aggregation

  • Hubo Guo,
  • Xin Li,
  • Qiang Wang,
  • Meiling Zhang,
  • Zhihong Huang

摘要

In the field of Person re-identification, text-guided semantic information mining has shown strong potential in recent years. The visual language model CLIP, with its excellent cross-modal comprehension capability, has achieved remarkable results in several tasks. In particular, in person re-identification, a two-stage CLIP-based approach has recently been proposed, which attempts to generate corresponding textual descriptions from the labels of specific pedestrians, which in turn aids the pedestrian classification task. However, after in-depth research, it was found that this approach has limitations in the first stage of feature extraction, i.e., the generated features are not sufficient to fully fit the pedestrian’s characteristics, which leads to the fact that the textual cues generated by utilizing these features in the second stage are not able to accurately represent the information for pedestrian in the vision. To solve this problem, this paper proposes an innovative Multi-stage Channel feature Aggregation method. The method digs deeper into the feature information of different stages through a well-designed feature aggregation strategy. This design allows our method to learn feature representations that are more attuned to pedestrian characteristics, which in turn generates more accurate textual cues. On the MSMT17, Market-1501, DukeMTMC, and Occluded-Duke datasets, we achieve accuracy improvements of 2.4%, 1.0%, 1.5%, and 3.6%, respectively. These results fully validate validity and applicability of our proposed method.