A Survey of HOI Detection: Method Evolution and Multimodal Fusion
摘要
As a key task in visual understanding, human object inter-action (HOI) detection aims to recognize people, objects and their inter-action relationships from static images. In recent years, with the rapid development of deep learning technology, HOI detection has made significant progress in method architecture, data resources and evaluation system. This paper systematically discusses and summarizes the past and the latest work of HOI detection, covering the mainstream evolution paths such as two-stage and single-stage methods, graph model and transformer architecture. In particular, this paper discusses the application of visual language model and large language model in HOI detection, including pre-trained feature transfer, text prompt driven zero sample detection, and cross modal knowledge injection. Finally, we summarize the current research trends and point out that weak supervised learning, open set recognition, complex interaction parsing and interpretable modeling are the key directions for future development. This paper aims to provide HOI researchers with a systematic, comprehensive and cutting-edge research map.