Click Video Object Segmentation (ClickVOS) targets the task of segmenting objects in videos with just one initial click on each object in the first frame. By transitioning from mask-based to click-based first-frame annotations, ClickVOS addresses the challenges of labor-intensive manual mask creation and the inflexibility in specifying arbitrary targets inherent in semi-supervised video object segmentation (Semi-VOS). However, previous ClickVOS methods typically lack granularity recognition capabilities, often resulting in ambiguous segmentation results. We present an innovative model termed “Granularity-ClickVOS” that allows for flexible selection of segmentation granularity. Specifically, we utilize a granularity-aware segmentation model that accepts click-generated point information to segment the target and propagates this mask across the remaining frames of the video. We validate our model on popular ClickVOS benchmarks and achieved astonishing results (70.1% J&F on DAVIS17-P).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Granularity-Aware Segment and Track for Click Video Object Segmentation

  • Chenhuan Liu,
  • Run Fang,
  • Chengxin Pang,
  • Xinhua Zeng

摘要

Click Video Object Segmentation (ClickVOS) targets the task of segmenting objects in videos with just one initial click on each object in the first frame. By transitioning from mask-based to click-based first-frame annotations, ClickVOS addresses the challenges of labor-intensive manual mask creation and the inflexibility in specifying arbitrary targets inherent in semi-supervised video object segmentation (Semi-VOS). However, previous ClickVOS methods typically lack granularity recognition capabilities, often resulting in ambiguous segmentation results. We present an innovative model termed “Granularity-ClickVOS” that allows for flexible selection of segmentation granularity. Specifically, we utilize a granularity-aware segmentation model that accepts click-generated point information to segment the target and propagates this mask across the remaining frames of the video. We validate our model on popular ClickVOS benchmarks and achieved astonishing results (70.1% J&F on DAVIS17-P).