Granularity-Aware Segment and Track for Click Video Object Segmentation
摘要
Click Video Object Segmentation (ClickVOS) targets the task of segmenting objects in videos with just one initial click on each object in the first frame. By transitioning from mask-based to click-based first-frame annotations, ClickVOS addresses the challenges of labor-intensive manual mask creation and the inflexibility in specifying arbitrary targets inherent in semi-supervised video object segmentation (Semi-VOS). However, previous ClickVOS methods typically lack granularity recognition capabilities, often resulting in ambiguous segmentation results. We present an innovative model termed “Granularity-ClickVOS” that allows for flexible selection of segmentation granularity. Specifically, we utilize a granularity-aware segmentation model that accepts click-generated point information to segment the target and propagates this mask across the remaining frames of the video. We validate our model on popular ClickVOS benchmarks and achieved astonishing results (70.1% J&F on DAVIS17-P).