<p>Recently, with the continuous development in the field of camouflaged object detection (COD), effectively separating objects highly similar to the background has become a focal point of research. Due to the high similarity between camouflaged objects and backgrounds, traditional single visual branch often perform poorly in such scenarios. To address this issue, we propose a multi-view learning detection network based on the Pyramid Vision Transformer, named Multi-view Learning for Camouflaged Object Detection with PVTv2 (MVLNet). By utilizing the information from RGB and noise views, our method can provide a more comprehensive description of the relationship between objects and backgrounds to improve the accuracy and robustness for COD. Inspired by human visual attention during observation, we design a Global Context Aggregation Module by using a U-shaped structure and progressively increasing dilation rates to simulate the human behavior of zooming in and out. Extensive experiments demonstrate that the proposed MVLNet outperforms 23 other representative models on three public datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-view learning for camouflaged object detection with PVTv2

  • Pu Yan,
  • Kang Ruan,
  • Lili Wang,
  • Yang Zhao,
  • Xu Wang

摘要

Recently, with the continuous development in the field of camouflaged object detection (COD), effectively separating objects highly similar to the background has become a focal point of research. Due to the high similarity between camouflaged objects and backgrounds, traditional single visual branch often perform poorly in such scenarios. To address this issue, we propose a multi-view learning detection network based on the Pyramid Vision Transformer, named Multi-view Learning for Camouflaged Object Detection with PVTv2 (MVLNet). By utilizing the information from RGB and noise views, our method can provide a more comprehensive description of the relationship between objects and backgrounds to improve the accuracy and robustness for COD. Inspired by human visual attention during observation, we design a Global Context Aggregation Module by using a U-shaped structure and progressively increasing dilation rates to simulate the human behavior of zooming in and out. Extensive experiments demonstrate that the proposed MVLNet outperforms 23 other representative models on three public datasets.