SPK: Semantic and Positional Knowledge for Zero-Shot Referring Expression Comprehension
摘要
Referring expression comprehension aims to localize an object in an image based on a natural language expression. This task is challenging due to the scarcity of large-scale annotated data, which prompts the research of zero-shot methods. In zero-shot scenarios, while existing models excel at grounding, they struggle to identify the target described by textual query due to the presence of multiple objects in a scene, as well as various spatial and attribute information. To address these issues, we propose a method called Semantic and Positional Knowledge (SPK), which leverages multimodal knowledge for fine-grained cross-modal matching in the referring expression comprehension task. Specifically, we pair words with visual representations as multimodal knowledge to match the information of expressions and images, such as objects, attributes, and spatial information. This method can be directly integrated with existing multimodal grounding models for further performance improvement. Experiments on the RefCOCO/+/g datasets demonstrate the effectiveness of our method, which can obtain consistent improvements.