Person re-identification is easily affected by various factors in unconstrained environments, and to improve its performance, it requires the support of auxiliary information. Although text descriptions of images provide detailed clues, the difficulty of modal fusion arises because images and text belong to different modalities. Therefore, to fully utilize and integrate the information from both image and text, we propose a new framework named Multimodal Feature Hierarchical Fusion(MFHF) that generates image captions automatically and fuses them with image features at the global level to improve person re-identification. Specifically, we design a dual-channel feature extraction network to learn global features from image and text information. Then, we propose a fusion strategy that combines multimodal features using a two-layer LSTM network. Meanwhile, we introduce the CMPM alignment function to reduce the feature gap between different modalities, ensuring consistency in multimodal feature representation and end-to-end training. Experimental validation on the CUHK-SYSU-Group dataset demonstrates the superiority of our proposed method over state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Feature Hierarchical Fusion for Text-Image Person Re-identification

  • Jiaxuan Li,
  • Likun Huang,
  • Chuanhu Zhu,
  • Song Zhang,
  • Qiang Li

摘要

Person re-identification is easily affected by various factors in unconstrained environments, and to improve its performance, it requires the support of auxiliary information. Although text descriptions of images provide detailed clues, the difficulty of modal fusion arises because images and text belong to different modalities. Therefore, to fully utilize and integrate the information from both image and text, we propose a new framework named Multimodal Feature Hierarchical Fusion(MFHF) that generates image captions automatically and fuses them with image features at the global level to improve person re-identification. Specifically, we design a dual-channel feature extraction network to learn global features from image and text information. Then, we propose a fusion strategy that combines multimodal features using a two-layer LSTM network. Meanwhile, we introduce the CMPM alignment function to reduce the feature gap between different modalities, ensuring consistency in multimodal feature representation and end-to-end training. Experimental validation on the CUHK-SYSU-Group dataset demonstrates the superiority of our proposed method over state-of-the-art approaches.