Fine-grained visual classification (FGVC) aims to identify subcategories that are visually highly similar yet differ in subtle ways, posing the significant challenge of small inter-class variance and large intra-class variance. We propose a novel method called Temperature-Guided Feature Extraction (TG-Net) to enhance the model’s sensitivity to discriminative regions and its robustness to background interference. TG-Net consists of two core modules: the Low-Temperature Extraction (LTE) module and the High-Temperature Refinement (HTR) module. The LTE module employs a temperature-controlled probabilistic sampling strategy to retain discriminative regions, enhances feature representation via a Semantic Information Enhancement (SIE) sub-module and graph convolutional network, and introduces a foreground-background contrastive learning mechanism to increase the separability between foreground and background features. The HTR module adopts a dual-path classification structure and incorporates a temperature modulation mechanism to jointly model semantic diversity and fine-grained local details. Experimental results on three widely-used FGVC benchmark datasets demonstrate that TG-Net achieves outstanding performance, exhibiting strong generalization and discriminative feature representation capabilities.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temperature-Guided Feature Extraction for Fine-Grained Visual Classification

  • Shixiong Wen,
  • Min Zhi

摘要

Fine-grained visual classification (FGVC) aims to identify subcategories that are visually highly similar yet differ in subtle ways, posing the significant challenge of small inter-class variance and large intra-class variance. We propose a novel method called Temperature-Guided Feature Extraction (TG-Net) to enhance the model’s sensitivity to discriminative regions and its robustness to background interference. TG-Net consists of two core modules: the Low-Temperature Extraction (LTE) module and the High-Temperature Refinement (HTR) module. The LTE module employs a temperature-controlled probabilistic sampling strategy to retain discriminative regions, enhances feature representation via a Semantic Information Enhancement (SIE) sub-module and graph convolutional network, and introduces a foreground-background contrastive learning mechanism to increase the separability between foreground and background features. The HTR module adopts a dual-path classification structure and incorporates a temperature modulation mechanism to jointly model semantic diversity and fine-grained local details. Experimental results on three widely-used FGVC benchmark datasets demonstrate that TG-Net achieves outstanding performance, exhibiting strong generalization and discriminative feature representation capabilities.