Multi-modal Knowledge-Enhanced Fine-Grained Image Classification
摘要
In image classification tasks, visual appearance is generally considered as a crucial cue for understanding images. However, relying solely on visual information can lead to misclassification in fine-grained image classification tasks. Multi-modal knowledge has been proven to provide critical cues for various computer vision tasks, such as image retrieval and vision question answering. In this paper, we integrate multi-modal knowledge into visual features to enhance the model’s understanding of visual content and accomplish fine-grained image classification tasks. Specifically, we adopt an effective visual enhanced module to capture global and local features, obtaining discriminative visual representations. Meanwhile, we employ knowledge distillation to transfer multi-modal knowledge from the Contrastive Language-Image Pre-training (CLIP) model to our model, improving its generalization ability. Moreover, we incorporate scene text into our visual features to provide richer contextual information. Experiments on the Con-Text, Drink Bottle, and Crowd Activity benchmark datasets demonstrate that our approach achieves 5.41%, 1.2%, and 7.55% improvements in mAP compared to the current state-of-the-art methods, respectively.