AF-DETR: efficient UAV small object detector via Assemble-and-Fusion mechanism
摘要
With the rise of deep learning networks, object detection technologies for unmanned aerial vehicle (UAV) have demonstrated outstanding performance in many application scenarios. However, current small object detection approaches overwhelmingly disregard sparse feature interactions and global context modeling, resulting in incomplete utilization and even loss of semantic information of small objects. Therefore, this study provides an advanced Assemble-and-Fusion mechanism used in DEtection TRansformer (AF-DETR), in which the aggregated global semantics are allocated across layers to augment fine-grained feature learning for small instances. Meanwhile, an adaptive context broadcasting module is designed to effectively integrate contextual information in the decoder, thus ensuring accurate detection of small objects. First, the last four stage features selected from the backbone are sent into the intra-scale feature interaction module, which performs self-attention operation on feature map of the last scale. Second, a fixed fusion module aligns and aggregates multi-scale representations prior to dissemination across layers. Features of adjoining levels then undergo transformation and consolidation within convolutional module. Finally, an enhanced adaptive context broadcasting module is introduced within the decoding MLP to incorporate aggregated semantics into individual tokens for broadcasting contextual information. Our AF-DETR achieves 49.5