<p>Fine-grained Visual Classification(FGVC) is challenging due to significant intra-class and subtle inter-class differences. Existing feature learning methods commonly rely on attention mechanisms to reinforce the discriminative features, where the seemingly contradictory feature suppression is absent. In this paper, we device a hierarchical bilateral architecture, and demonstrate combining feature refinement with feature suppression benefits FGVC clearly. By suppressing the most salient regions, richer complementary semantic regions that are discriminative for visually similar subclasses are detected. On the other hand, features from the regions of interest are reinforced by a weighted map specifying hierarchical cross-stage attentions, spontaneously highlighting the dominant visual cues with different granularities. Alternating between feature suppression and feature reinforcement, the network gradually learns complementary discriminative features from coarse-grained to fine-grained. Extensive experiments conducted on three benchmarked FGVC datasets show that the proposed method achieves competitive results with the state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature reinforcement meets feature suppression: a hierarchical bilateral method for fine-grained visual classification

  • Cheng Pang,
  • Yingjie Song,
  • Dingzhou Xie,
  • Rushi Lan

摘要

Fine-grained Visual Classification(FGVC) is challenging due to significant intra-class and subtle inter-class differences. Existing feature learning methods commonly rely on attention mechanisms to reinforce the discriminative features, where the seemingly contradictory feature suppression is absent. In this paper, we device a hierarchical bilateral architecture, and demonstrate combining feature refinement with feature suppression benefits FGVC clearly. By suppressing the most salient regions, richer complementary semantic regions that are discriminative for visually similar subclasses are detected. On the other hand, features from the regions of interest are reinforced by a weighted map specifying hierarchical cross-stage attentions, spontaneously highlighting the dominant visual cues with different granularities. Alternating between feature suppression and feature reinforcement, the network gradually learns complementary discriminative features from coarse-grained to fine-grained. Extensive experiments conducted on three benchmarked FGVC datasets show that the proposed method achieves competitive results with the state-of-the-art methods.