Feature fusion means a lot to DETRs
摘要
The DEtection TRansformer (DETR) is a pioneering end-to-end object detector that has garnered significant success. However, it is not without its shortcomings, as it is characterized by a slow convergence rate and requires substantial computational resources. Subsequent research efforts, such as Deformable DETR, have been directed at enhancing the performance of DETR and accelerating its convergence process. Nonetheless, in adopting a novel architecture, DETR forgoes some of the well-established designs that have been integral to the success of the YOLO series. Specifically, the multi-scale feature maps and feature fusion are key elements contributing to the success of the YOLO series. In contrast, DETR models do not fully leverage these mature designs, leading to less feature extraction compared to YOLO. To harness the benefits of feature fusion in DETR models and thereby extract richer features, we introduce the Feature-Fusion DETR (FF-DETR), which combines the advantages of both DETR models and the YOLO series. Our experiments demonstrate that with meticulous design, our model can surpass the performance of existing DETR models. For instance, under a training regime of just 12 epochs, FF-DETR augments the performance of DINO (ResNet-50) by a notable +0.3% in Average Precision (AP), achieving an impressive 49.1% AP on the COCO 2017 datasets. And under 24 epochs, FF-DETR achieves 50.9% AP on the COCO 2017 datasets. Code will be available at https://github.com/huakaixu/FF-DETR