<p>Individual cattle identification provides excellent and fundamental technology for modern animal husbandry. Most of the existing cattle identification methods use the RGB modality, and this single modality provides minimal feature information. In addition, RGB data is weakly robust in environments with uneven illumination and substantial background interference, resulting in a significant decrease in recognition accuracy. To address the challenges of cattle identification, we propose a novel dual-stream modality prompt network (2S-MPN), which integrates image data and text descriptions to achieve complementary feature learning and achieves excellent performance. The 2S-MPN consists of an RGB stream, an Infrared stream, a modality prompt module, and a feature-weighted fusion module. The RGB stream uses a convolutional neural network to efficiently extract the biometric features of a cattle face or muzzle pattern. The Infrared stream uses a lightweight transformer network to establish long-range context dependencies for infrared images. The interaction of RGB and Infrared stream effectively establishes comprehensive local and global feature information for the network. Inspired by prompt learning, we designed the modality prompt module, which introduces textual feature information to the network and provides guidance for RGB and infrared images. We design the feature-weighted fusion module, which generates weight vectors from similarity metrics to fuse the three modal feature information more effectively. Experimental results on two homemade datasets show that the 2S-MPN is competitive in the cattle identification task.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-stream modality prompt network for individual cattle identification

  • Mei Yang,
  • Qi Li,
  • Jianmin Zhao,
  • Yueming Wang,
  • Jingfang Gao,
  • Fangfang Xue,
  • Dongxu Li

摘要

Individual cattle identification provides excellent and fundamental technology for modern animal husbandry. Most of the existing cattle identification methods use the RGB modality, and this single modality provides minimal feature information. In addition, RGB data is weakly robust in environments with uneven illumination and substantial background interference, resulting in a significant decrease in recognition accuracy. To address the challenges of cattle identification, we propose a novel dual-stream modality prompt network (2S-MPN), which integrates image data and text descriptions to achieve complementary feature learning and achieves excellent performance. The 2S-MPN consists of an RGB stream, an Infrared stream, a modality prompt module, and a feature-weighted fusion module. The RGB stream uses a convolutional neural network to efficiently extract the biometric features of a cattle face or muzzle pattern. The Infrared stream uses a lightweight transformer network to establish long-range context dependencies for infrared images. The interaction of RGB and Infrared stream effectively establishes comprehensive local and global feature information for the network. Inspired by prompt learning, we designed the modality prompt module, which introduces textual feature information to the network and provides guidance for RGB and infrared images. We design the feature-weighted fusion module, which generates weight vectors from similarity metrics to fuse the three modal feature information more effectively. Experimental results on two homemade datasets show that the 2S-MPN is competitive in the cattle identification task.