<p>Human object interaction (HOI) is a core visual task in computer vision for understanding the static relationship between humans and objects. Recently, some approaches have achieved impressive results by Vision-Language Models (VLM) to provide prior knowledge for HOI detectors. However, such methods often fail to effectively extract knowledge features. In this paper, we propose a Multi-Query Network (MQN) to address the drawbacks of the prior query-based HOI detectors in two dimensions. For the breadth, the previous method assigning very few queries to a single instance reduces training efficiency. We design a multi-query branch that allows one instance to correspond to multiple queries while retaining the original query scheme. This approach greatly enhances knowledge transfer capabilities. For the depth, to address the cascading issues of decoding sequences, we introduce a Knowledge Representation Extractor (KRE) that accumulates intermediate features progressively through the decoding layers. Furthermore, we propose a Verb Attention (VA) module to enhance supervision over verb categories. Extensive experiments have shown that MQN is significantly more effective than the state of the art on multiple datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-query network learning for hoi detection via vision-language models

  • Leideng Shi,
  • Juan Zhang

摘要

Human object interaction (HOI) is a core visual task in computer vision for understanding the static relationship between humans and objects. Recently, some approaches have achieved impressive results by Vision-Language Models (VLM) to provide prior knowledge for HOI detectors. However, such methods often fail to effectively extract knowledge features. In this paper, we propose a Multi-Query Network (MQN) to address the drawbacks of the prior query-based HOI detectors in two dimensions. For the breadth, the previous method assigning very few queries to a single instance reduces training efficiency. We design a multi-query branch that allows one instance to correspond to multiple queries while retaining the original query scheme. This approach greatly enhances knowledge transfer capabilities. For the depth, to address the cascading issues of decoding sequences, we introduce a Knowledge Representation Extractor (KRE) that accumulates intermediate features progressively through the decoding layers. Furthermore, we propose a Verb Attention (VA) module to enhance supervision over verb categories. Extensive experiments have shown that MQN is significantly more effective than the state of the art on multiple datasets.