In recent years, infrared video object segmentation has found extensive applications across various domains, including military surveillance, fire rescue and other fields. Despite the considerable potential of infrared imaging, segmenting objects in infrared video sequences remains a challenging task due to factors such as numerous video frames, limited available data, and complex backgrounds. To address these challenges, we propose a robust solution employing a large model adaptation strategy tailored for infrared datasets, coupled with a one-shot training approach to leverage information across video frames. Our framework, built upon the Segment Anything Model (SAM), effectively extends the parameters of a large model to accommodate infrared images, bridging the gap between training and testing video data and enhancing segmentation performance. The methodology involves supervised and unsupervised training segments, utilizing consistent and contrast loss mechanisms to ensure the model’s robustness and accuracy. Our approach has demonstrated experimentally its capability to effectively migrate parameters trained on visible light to the infrared domain, yielding excellent performance on the infrared dataset VTUAV.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Source-Free One-Shot Infrared Video Object Segmentation Based on the Segment Anything Model

  • Jiaqi Chen,
  • Dingwen Zhang,
  • Weinan Zhao,
  • Lei Li,
  • Jun Ren,
  • Hang Qi,
  • Ruitao Lu,
  • Junwei Han

摘要

In recent years, infrared video object segmentation has found extensive applications across various domains, including military surveillance, fire rescue and other fields. Despite the considerable potential of infrared imaging, segmenting objects in infrared video sequences remains a challenging task due to factors such as numerous video frames, limited available data, and complex backgrounds. To address these challenges, we propose a robust solution employing a large model adaptation strategy tailored for infrared datasets, coupled with a one-shot training approach to leverage information across video frames. Our framework, built upon the Segment Anything Model (SAM), effectively extends the parameters of a large model to accommodate infrared images, bridging the gap between training and testing video data and enhancing segmentation performance. The methodology involves supervised and unsupervised training segments, utilizing consistent and contrast loss mechanisms to ensure the model’s robustness and accuracy. Our approach has demonstrated experimentally its capability to effectively migrate parameters trained on visible light to the infrared domain, yielding excellent performance on the infrared dataset VTUAV.