<p>Person re-identification (Re-ID) is a technique employed to recognize a target pedestrian based on source information and video sequences from non-overlapping camera fields of view. This technique has garnered increasing attention due to its extensive application prospects in intelligent video surveillance within the internet of things and other contexts. In this paper, we propose a multi-task learning approach that incorporates text information to enhance recognition accuracy. We employ text as auxiliary data, utilizing a dual-stream transformer’s encoder to extract both image and text features. To further improve the model’s interaction and feature learning capabilities, we introduce a cross-modal interaction encoder (CIE) and a feature sharing learning (FSL) network. The CIE facilitates fine-grained alignment between text and image features, while the FSL network learns modality-invariant feature representations. Our method, DSFAT, demonstrates superior performance in person re-identification tasks when compared to state-of-the-art methods, as validated on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSFAT: a dual-stream framework assisted by textual information for person re-identification in real scenes

  • Xuanrui Xiong,
  • Haihong Huang,
  • Tianyu Li,
  • Xiaolin Fan,
  • Yuan Zhang

摘要

Person re-identification (Re-ID) is a technique employed to recognize a target pedestrian based on source information and video sequences from non-overlapping camera fields of view. This technique has garnered increasing attention due to its extensive application prospects in intelligent video surveillance within the internet of things and other contexts. In this paper, we propose a multi-task learning approach that incorporates text information to enhance recognition accuracy. We employ text as auxiliary data, utilizing a dual-stream transformer’s encoder to extract both image and text features. To further improve the model’s interaction and feature learning capabilities, we introduce a cross-modal interaction encoder (CIE) and a feature sharing learning (FSL) network. The CIE facilitates fine-grained alignment between text and image features, while the FSL network learns modality-invariant feature representations. Our method, DSFAT, demonstrates superior performance in person re-identification tasks when compared to state-of-the-art methods, as validated on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets.