Understanding scenes requires not only the detection objects but also the recognition of the interactions between them. Human-Object Interaction (HOI) detection plays a crucial role in enhancing contextual comprehension by identifying the interactions between humans and objects, which is essential for building more robust and intelligent vision systems. While DETR-based models have shown significant success in HOI detection, they are hindered by slow training convergence. The SOV-STG method has attempted to address this challenge in previous research. To further improve the learning efficiency and accuracy of SOV-STG, we introduce a novel Separate Guided Denoising training strategy specifically designed for HOI detection. Our approach separates the denoising of noised ground truth data for both the human-object decoder and the verb decoder, enabling more efficient and targeted training. Furthermore, we enhance training performance by merging redundant human-object pair annotations, and filtering and regenerating noised bounding boxes. The proposed method was validated on the HICO-DET dataset, achieving state-of-the-art results. Our contributions include a novel training strategy that improves accuracy and ablation studies demonstrating its effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Separate Guided Denoising Training for Human-Object Interaction Detection

  • Yuki Isoda,
  • Daisuke Kobayashi

摘要

Understanding scenes requires not only the detection objects but also the recognition of the interactions between them. Human-Object Interaction (HOI) detection plays a crucial role in enhancing contextual comprehension by identifying the interactions between humans and objects, which is essential for building more robust and intelligent vision systems. While DETR-based models have shown significant success in HOI detection, they are hindered by slow training convergence. The SOV-STG method has attempted to address this challenge in previous research. To further improve the learning efficiency and accuracy of SOV-STG, we introduce a novel Separate Guided Denoising training strategy specifically designed for HOI detection. Our approach separates the denoising of noised ground truth data for both the human-object decoder and the verb decoder, enabling more efficient and targeted training. Furthermore, we enhance training performance by merging redundant human-object pair annotations, and filtering and regenerating noised bounding boxes. The proposed method was validated on the HICO-DET dataset, achieving state-of-the-art results. Our contributions include a novel training strategy that improves accuracy and ablation studies demonstrating its effectiveness.