<p>Camouflaged object detection (COD) aims to identify objects that are visually indistinguishable from their surroundings. Existing methods typically enhance detection by introducing spatial-domain and boundary-aware features, striving to increase the discriminability of spatial representations to suppress background interference. However, spatial-domain features often exhibit limited sensitivity to boundary cues and the subtle structures of camouflaged objects. To address this issue, we propose a multi-scale fusion framework that combines spatial and frequency-domain information for robust COD. Specifically, we enhance semantic representations of camouflaged objects by integrating multi-scale spatial features with frequency-domain components. A spatial-frequency perception mechanism is designed to suppress background noise across different scales, while a reverse attention strategy is employed for progressive decoding in both spatial and frequency domains. Extensive experiments are conducted on three benchmark datasets, comparing our approach with 14 spatial-domain-based and 4 frequency-domain-based methods, and the results demonstrate that our method achieves the state-of-the-art.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale spatial-frequency fusion for camouflaged object detection

  • Yuqiao Song,
  • Benyi Zhang,
  • Peng Li

摘要

Camouflaged object detection (COD) aims to identify objects that are visually indistinguishable from their surroundings. Existing methods typically enhance detection by introducing spatial-domain and boundary-aware features, striving to increase the discriminability of spatial representations to suppress background interference. However, spatial-domain features often exhibit limited sensitivity to boundary cues and the subtle structures of camouflaged objects. To address this issue, we propose a multi-scale fusion framework that combines spatial and frequency-domain information for robust COD. Specifically, we enhance semantic representations of camouflaged objects by integrating multi-scale spatial features with frequency-domain components. A spatial-frequency perception mechanism is designed to suppress background noise across different scales, while a reverse attention strategy is employed for progressive decoding in both spatial and frequency domains. Extensive experiments are conducted on three benchmark datasets, comparing our approach with 14 spatial-domain-based and 4 frequency-domain-based methods, and the results demonstrate that our method achieves the state-of-the-art.