Accurate delineation of boundaries and instance semantics is crucial for tasks like object localization in robotic arm grasping, and vehicle and pedestrian detection in autonomous driving. While research often focuses on improving instance segmentation accuracy and lightweight models, the importance of boundary detection and open-vocabulary capabilities for human-level perception is often overlooked. In this work, we propose a lightweight visual-language dual-task framework, IS-Goal, that simultaneously performs instance segmentation and boundary detection under open-vocabulary. It includes a prompt text encoder, a two-stream image encoder, and a visual-language adaptive weight decoder (VL-AWD) for multi-level cross-modal feature fusion. The text encoder extracts text embeddings, the two-stream image encoder captures instance and boundary features, and the VL-AWD module learns channel relationships to obtain adaptive weight allocation for instance features and instance boundary features, enabling multi-modal fusion. Additionally, we introduce a regularization loss to mitigate the conflicts in dual-task learning and diverse deep supervision. Compared to existing methods, IS-Goal improves instance segmentation and boundary detection performance under open-vocabulary. We first validate IS-Goal’s effectiveness in open-vocabulary instance segmentation tasks on the MS COCO dataset for identifying and distinguishing new categories from base categories. Subsequently, on the LVIS dataset, IS-Goal surpasses existing dual-task methods with a boundary AP of 27.5%, instance segmentation AP of 37.3%, and ODS/OIS scores of 67.7/68.2. Zero-shot performance on PASCAL VOC2012 is demonstrated with an inference speed of 15.7 FPS on an RTX 2080 Ti GPU with 500 × 500 input resolution.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Open-Vocabulary Instance Segmentation-Boundary IS-Goal

  • Quan Tang

摘要

Accurate delineation of boundaries and instance semantics is crucial for tasks like object localization in robotic arm grasping, and vehicle and pedestrian detection in autonomous driving. While research often focuses on improving instance segmentation accuracy and lightweight models, the importance of boundary detection and open-vocabulary capabilities for human-level perception is often overlooked. In this work, we propose a lightweight visual-language dual-task framework, IS-Goal, that simultaneously performs instance segmentation and boundary detection under open-vocabulary. It includes a prompt text encoder, a two-stream image encoder, and a visual-language adaptive weight decoder (VL-AWD) for multi-level cross-modal feature fusion. The text encoder extracts text embeddings, the two-stream image encoder captures instance and boundary features, and the VL-AWD module learns channel relationships to obtain adaptive weight allocation for instance features and instance boundary features, enabling multi-modal fusion. Additionally, we introduce a regularization loss to mitigate the conflicts in dual-task learning and diverse deep supervision. Compared to existing methods, IS-Goal improves instance segmentation and boundary detection performance under open-vocabulary. We first validate IS-Goal’s effectiveness in open-vocabulary instance segmentation tasks on the MS COCO dataset for identifying and distinguishing new categories from base categories. Subsequently, on the LVIS dataset, IS-Goal surpasses existing dual-task methods with a boundary AP of 27.5%, instance segmentation AP of 37.3%, and ODS/OIS scores of 67.7/68.2. Zero-shot performance on PASCAL VOC2012 is demonstrated with an inference speed of 15.7 FPS on an RTX 2080 Ti GPU with 500 × 500 input resolution.