<p>To address the issues of structural information loss and object occlusion arising from existing camouflaged object detection methods in handling complex situations, we propose a novel network that integrates scale awareness and enhanced large kernel attention (SALK-Net). Specifically, our network takes ternary images as input to mine the additional information contained at different scales. Firstly, we use a shared feature encoder to extract features and align channels from multi-scale input images. Secondly, enhanced large kernel attention is introduced to guide the fusion of scale features, which aims to fully perceive global semantic information and minimize the loss of valuable clues. Thirdly, in the designed hybrid-scale mixed-scale decoder, we adopt a progressive structure to explore and gradually accumulate the clue information contained in the feature channels. Finally, a dynamic weighting strategy for boundary and structure is introduced to loss constraints together with prior knowledge to help the model predict challenging pixels. We compared the proposed model with 12 state-of-the-art methods in 4 public datasets. The results were then assessed on 4 metrics. The structural similarity measure and enhanced alignment measure in a large trained dataset reached 0.861 and 0.927 respectively whereas 0.872 and 0.926 respectively in the untrained large dataset, which demonstrates the competitiveness of our method over state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A ternary encoding network fusing scale awareness and large kernel attention for camouflaged object detection

  • Chaoquan Zheng,
  • Jinzheng Lu,
  • Kun Hu,
  • Qiang Xiang,
  • Ling Miao

摘要

To address the issues of structural information loss and object occlusion arising from existing camouflaged object detection methods in handling complex situations, we propose a novel network that integrates scale awareness and enhanced large kernel attention (SALK-Net). Specifically, our network takes ternary images as input to mine the additional information contained at different scales. Firstly, we use a shared feature encoder to extract features and align channels from multi-scale input images. Secondly, enhanced large kernel attention is introduced to guide the fusion of scale features, which aims to fully perceive global semantic information and minimize the loss of valuable clues. Thirdly, in the designed hybrid-scale mixed-scale decoder, we adopt a progressive structure to explore and gradually accumulate the clue information contained in the feature channels. Finally, a dynamic weighting strategy for boundary and structure is introduced to loss constraints together with prior knowledge to help the model predict challenging pixels. We compared the proposed model with 12 state-of-the-art methods in 4 public datasets. The results were then assessed on 4 metrics. The structural similarity measure and enhanced alignment measure in a large trained dataset reached 0.861 and 0.927 respectively whereas 0.872 and 0.926 respectively in the untrained large dataset, which demonstrates the competitiveness of our method over state-of-the-art methods.