Enhancing few-shot object detection via visual–semantic calibration and dual-localization refinement
摘要
Few-shot object detection (FSOD) aims to accurately detect novel object classes using limited annotated samples, crucial for applications like medical diagnosis and rare species monitoring. Traditional approaches often rely on transfer learning, yet they face challenges in feature representation calibration and region proposal localization. To address these, we introduce Visual–Semantic Joint Calibration and Localization Enhancement (VSJCLE), a fine-tuning framework that enhances feature representation and region proposal localization. VSJCLE incorporates a Visual–Semantic Joint Calibration (VSJC) module, which dynamically refines image features using both visual and semantic information, significantly improving feature discriminability. Additionally, the Twofold Localization Enhancement (TLE) strategy, comprising Dual-RPN and Hierarchical Feature Enhancement (HFE) submodules, achieves a balance between stability and adaptability, enhancing detection localization. Experimental results on the PASCAL VOC and MS COCO datasets demonstrate that VSJCLE consistently outperforms state-of-the-art methods in both FSOD and Generalized FSOD (G-FSOD) settings. Notably, on MS COCO under the 30-shot setting, VSJCLE achieves an AP of 32.3%, surpassing the second-best model by 8.7%. These results suggest that VSJCLE exhibits strong generalization on standard benchmarks and has potential relevance to practical detection tasks.