<p>With the rapid development of intelligent security systems, the demand for vehicle re-identification has surged exponentially. Vehicle re-identification involves recognizing the same vehicle across different camera perspectives, necessitating robust local feature processing. While transformers have shown promising results in this field, their inherent self-attention mechanism tends to dilute high-frequency texture details, hindering local feature extraction. Additionally, challenges such as occlusion and misalignment can lead to information loss and noise introduction, reducing re-identification accuracy. To address these issues, we introduce the frequency transformer with local feature enhancement (LFFT). The proposed framework comprises a frequency layer and a jigsaw select patches module (JSPM). The frequency layer enhances the weights of high-frequency component features using fast Fourier transform to improve local feature extraction at the lower layers. Meanwhile, the attention layer at the higher layers continues to extract global features. The JSPM incorporates discriminative patches obtained from attention layers into randomly shuffled and reorganized groups, enhancing the global discriminative capability of local features. The method does not utilize additional information or auxiliary networks. Experimental evaluations on two vehicle re-identification datasets, VeRi-776 and VehicleID, demonstrate the effectiveness of our method compared to recent approaches. The code is available at <a href="https://github.com/xianghlin/LFFT">https://github.com/xianghlin/LFFT</a>, accompanied by detailed usage instructions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frequency transformer with local feature enhancement for improved vehicle re-identification

  • Honglin Xiang,
  • Jiahao Wang,
  • Yulong Sun,
  • Ming Ye

摘要

With the rapid development of intelligent security systems, the demand for vehicle re-identification has surged exponentially. Vehicle re-identification involves recognizing the same vehicle across different camera perspectives, necessitating robust local feature processing. While transformers have shown promising results in this field, their inherent self-attention mechanism tends to dilute high-frequency texture details, hindering local feature extraction. Additionally, challenges such as occlusion and misalignment can lead to information loss and noise introduction, reducing re-identification accuracy. To address these issues, we introduce the frequency transformer with local feature enhancement (LFFT). The proposed framework comprises a frequency layer and a jigsaw select patches module (JSPM). The frequency layer enhances the weights of high-frequency component features using fast Fourier transform to improve local feature extraction at the lower layers. Meanwhile, the attention layer at the higher layers continues to extract global features. The JSPM incorporates discriminative patches obtained from attention layers into randomly shuffled and reorganized groups, enhancing the global discriminative capability of local features. The method does not utilize additional information or auxiliary networks. Experimental evaluations on two vehicle re-identification datasets, VeRi-776 and VehicleID, demonstrate the effectiveness of our method compared to recent approaches. The code is available at https://github.com/xianghlin/LFFT, accompanied by detailed usage instructions.