Segment Anything (SAM) has become popular in segmentation, yielding competitive results because it can accurately segment almost all objects in natural images and quickly adapt to downstream tasks when combined with fine-tuning. However, we find that simply finetuning can degrade the model’s generalization on out-of-domain data. This paper proposes TP-SAM to improve generalization, enhancing performance metrics on downstream and out-of-domain datasets. The fundamental idea is to have learnable parameters for different types of prompts, thereby decoupling three distinct tasks, and optimizing the model’s loss on masks generated through shared robust output token (ROT) across tasks. Particularly, to ensure sufficient interac-tion between newly added parameters and image features, etc., image features and tokens are repeatedly fed into the mask decoder multiple times. During inference, a combination of the ROT token and task-specific prompts is fed into the mask decoder for each task. Extensive experiments on several segmentation datasets under different types of prompts demonstrate that our method not only improves metrics within the domain but also enhances metrics for zero-shot segmentation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TP-SAM: Fine-Tuning SAM with Task-Specific Prompt in the Loop

  • Tingting Xiao,
  • Yinzhou Ling

摘要

Segment Anything (SAM) has become popular in segmentation, yielding competitive results because it can accurately segment almost all objects in natural images and quickly adapt to downstream tasks when combined with fine-tuning. However, we find that simply finetuning can degrade the model’s generalization on out-of-domain data. This paper proposes TP-SAM to improve generalization, enhancing performance metrics on downstream and out-of-domain datasets. The fundamental idea is to have learnable parameters for different types of prompts, thereby decoupling three distinct tasks, and optimizing the model’s loss on masks generated through shared robust output token (ROT) across tasks. Particularly, to ensure sufficient interac-tion between newly added parameters and image features, etc., image features and tokens are repeatedly fed into the mask decoder multiple times. During inference, a combination of the ROT token and task-specific prompts is fed into the mask decoder for each task. Extensive experiments on several segmentation datasets under different types of prompts demonstrate that our method not only improves metrics within the domain but also enhances metrics for zero-shot segmentation.