<p>To address the challenges of inadequate extraction of features from interaction regions due to occlusion and the limited generalization in small-sample training, we propose a human-object interaction detection algorithm based on visual-semantic interaction perception. To improve the extraction of interaction region features hindered by occlusion, we propose a Deep Visual Interaction Feature Extraction Network(DVIF). Using graph convolution to develop an adaptable human-object relationship graph, paired with an optimized human pose segment, allows us to better capture significant interaction region characteristics. To enhance the generalization ability of small sample training data, a Semantic-driven Sample Expansion Module (SSE) is proposed. This module incorporates word embeddings and linguistic priors into data generation to augment the training samples. We propose a Cross-modal External Attention Fusion Module (CEA) to synergistically enhance visual and semantic information, thereby improving the accuracy and robustness of interaction detection. Experiments on the HICO-DET and V-COCO datasets achieved human-object interaction detection accuracy of 30.67% and 60.3%, demonstrating the effectiveness of the proposed algorithm.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human-Object interaction detection algorithm based on visual-semantic interaction perception

  • Qing Ye,
  • Gaohui Deng,
  • Xiuju Xu,
  • Yongmei Zhang

摘要

To address the challenges of inadequate extraction of features from interaction regions due to occlusion and the limited generalization in small-sample training, we propose a human-object interaction detection algorithm based on visual-semantic interaction perception. To improve the extraction of interaction region features hindered by occlusion, we propose a Deep Visual Interaction Feature Extraction Network(DVIF). Using graph convolution to develop an adaptable human-object relationship graph, paired with an optimized human pose segment, allows us to better capture significant interaction region characteristics. To enhance the generalization ability of small sample training data, a Semantic-driven Sample Expansion Module (SSE) is proposed. This module incorporates word embeddings and linguistic priors into data generation to augment the training samples. We propose a Cross-modal External Attention Fusion Module (CEA) to synergistically enhance visual and semantic information, thereby improving the accuracy and robustness of interaction detection. Experiments on the HICO-DET and V-COCO datasets achieved human-object interaction detection accuracy of 30.67% and 60.3%, demonstrating the effectiveness of the proposed algorithm.