Foreground-Centric learning improves robustness in fine-grained visual recognition
摘要
Fine-grained visual recognition—distinguishing categories that differ by subtle, localized cues—often degrades in the wild when background clutter and pose variation dominate the signal. We adopt a simple principle: isolate the subject before recognition. By constructing subject-centric views and training on these foregrounds instead of full scenes, decisions are steered toward diagnostic morphology rather than incidental context. Evaluated on a challenging ecological case study of visually similar species, the approach yields consistent gains in accuracy and per-class reliability, with the largest improvements for confusable categories. The procedure is modular, annotation-light, and compatible with diverse learning architectures and deployment settings. We release open resources and a reproducibility checklist to enable transparent evaluation; source code and supporting data are available at https://github.com/AliAlfatemi/warbler_yolo. Beyond biodiversity monitoring, foreground-centric learning offers a broadly applicable path to more reliable fine-grained recognition in agriculture, product inspection, and biomedical imaging.