RGB-T crowd counting integrates visible (RGB) and thermal (T) images to estimate crowd density. This task faces two critical challenges: modality misalignment and weak feature representation. To tackle these challenges, we present an innovative approach by introducing a hybrid loss function with enhanced cross-modal features as feedback. Our approach comprises two key components: a feature enhancement module that iteratively refines and amplifies key features from both modalities and a semantic alignment hybrid loss, which progressively aligns RGB and thermal representations across network hierarchies. Extensive experiments on RGB-T crowd counting benchmarks substantiate the effectiveness of our approach, achieving an average improvement in accuracy of 6% on RGBT-CC and 14% on DroneRGBT compared to previous methods, reaching state-of-the-art performance in RGB-T crowd counting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Hybrid Loss with Enhanced Feature Feedback for RGB-T Crowd Counting

  • Yaodan Yu,
  • Zengxu Liang

摘要

RGB-T crowd counting integrates visible (RGB) and thermal (T) images to estimate crowd density. This task faces two critical challenges: modality misalignment and weak feature representation. To tackle these challenges, we present an innovative approach by introducing a hybrid loss function with enhanced cross-modal features as feedback. Our approach comprises two key components: a feature enhancement module that iteratively refines and amplifies key features from both modalities and a semantic alignment hybrid loss, which progressively aligns RGB and thermal representations across network hierarchies. Extensive experiments on RGB-T crowd counting benchmarks substantiate the effectiveness of our approach, achieving an average improvement in accuracy of 6% on RGBT-CC and 14% on DroneRGBT compared to previous methods, reaching state-of-the-art performance in RGB-T crowd counting.