<p>As an intersection of computer vision and natural language understanding, language-guided video segmentation represents an important research field by enabling precise pixel-level object identification through text query instructions. This survey tracks the progression of language-guided video segmentation from Referring Video Object Segmentation (RVOS) to the emerging Reasoning Video Object Segmentation (ReasonVOS). Specifically, RVOS focuses on grounding explicit text queries to visual content, while ReasonVOS advances it by interpreting implicit queries that demand world knowledge and multi-step reasoning for object identification, where we introduce a taxonomy to organize these existing methods across both paradigms. We also provide an analysis of the corresponding benchmark datasets, alongside evaluation methods that extend beyond traditional segmentation metrics to assess reasoning correctness. Finally, we provide a discussion for the applications of language-guided video segmentation and their existing challenges to inspire future exploration. Generally, this survey aims to serve as both a reference for researchers entering this field and a roadmap for advancing language-guided video understanding toward more intelligent and practical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A survey of language-guided video object segmentation: from referring to reasoning

  • Yiqing Shen,
  • Dell Zhang

摘要

As an intersection of computer vision and natural language understanding, language-guided video segmentation represents an important research field by enabling precise pixel-level object identification through text query instructions. This survey tracks the progression of language-guided video segmentation from Referring Video Object Segmentation (RVOS) to the emerging Reasoning Video Object Segmentation (ReasonVOS). Specifically, RVOS focuses on grounding explicit text queries to visual content, while ReasonVOS advances it by interpreting implicit queries that demand world knowledge and multi-step reasoning for object identification, where we introduce a taxonomy to organize these existing methods across both paradigms. We also provide an analysis of the corresponding benchmark datasets, alongside evaluation methods that extend beyond traditional segmentation metrics to assess reasoning correctness. Finally, we provide a discussion for the applications of language-guided video segmentation and their existing challenges to inspire future exploration. Generally, this survey aims to serve as both a reference for researchers entering this field and a roadmap for advancing language-guided video understanding toward more intelligent and practical applications.