Multi-attention Mechanism for Few-Shot Object Detection
摘要
The majority of existing few-shot object detection techniques depend on conventional deep learning architectures and can be categorized into either single-stage or two-stage approaches, both of which have demonstrated corresponding outcomes. In the context of two-stage object detection models, however, the challenge often lies in the generation of numerous irrelevant bounding boxes, largely due to the limited number of sample instances. To tackle this very problem, the current study presents a novel few-shot object detection algorithm christened Multi-attention for Few-Shot Object Detection (MuA-FSOD), which ingeniously integrates multiple attention mechanisms. This algorithm enhances the feature representation of new classes through feature fusion techniques and alleviates the adverse effects that may be introduced by feature fusion through the Cross Attention Redistributed (CAReD) attention mechanism, while maintaining the richness of features. Additionally, the algorithm employs the Query-Support Attention Module (QSAM) attention mechanism in the relationship matching phase, improving the quality of candidate boxes by assessing the mutual relationship within the interaction between support and query images. The experimental outcomes affirm that the proposed model exceeds the performance of the baseline Few-Shot Object Detection (FSOD) method on the MS COCO benchmark across the \(\text{AP}\) , \({\text{AP}}_{50}\) , \({\text{AP}}_{75}\) metrics, with improvements of 0.9%, 1.2%, and 0.8% respectively. Additionally, it delivers commendable results when tested under the N-way K-shot setting on the PASCAL VOC dataset. These results collectively demonstrate that this method significantly uplifts the detection accuracy for scenarios involving few-shot object recognition tasks.