LMNet:A location Feature Generation and Memory Token Interactive Network for camouflaged object detection
摘要
Camouflaged Object Detection (COD) requires models to accurately detect objects that are seamlessly concealed within their surroundings. Existing COD methods primarily focus on extracting camouflaged cues from contextual information, often ignoring the latent information in background regions, which can provide beneficial guidance for the models to detect the edges of camouflaged objects. However, since background regions are usually very complex, it is challenging to extract the latent information from them. To address this challenge, we propose a novel location feature generation and memory token interaction network (LMNet) for COD. This network employs the Swin Transformer as its backbone and integrates three functional modules designed by us, i.e., the location information enhancement module, the memory token interaction module and the triple knowledge aggregation module. In our model, we first enhance the location information of camouflaged objects by the location information enhancement modules, and fully mine the information of camouflaged objects and the latent information of background regions by the memory token interaction modules. Then, we uses the triple knowledge aggregation modules to fuse the features output by the two types of modules and the features extracted by the backbone, obtaining the fused features with enhanced camouflaged object information. Experimental results on four benchmark datasets confirm that our network exhibits strong generalization capability and performs better than other COD methods.