<p>Current few-shot fine-grained image classification models primarily rely on mining discriminative local features to enhance inter-class separability. However, these models remain insufficient in effectively addressing the challenge of substantial intra-class variation. To address the aforementioned challenges, this paper proposes a feature disentanglement and reconstruction network based on autoencoders. This network employs the disentanglement design with content and style encoders to separately process spatial information and high-level semantic information in images. It utilizes a decoder for feature reconstruction, which can effectively achieve cross-image feature alignment. By reconstructing the features, the model’s ability to represent intra-class variability in few-shot fine-grained classification tasks has been enhanced. The experimental results demonstrate that the proposed model significantly improves classification accuracy on three canonical few-shot datasets (Stanford Dogs, Stanford Cars, and CUB-200–2011). Specifically, the model achieves a 4.08% accuracy improvement in 1-shot classification tasks on the Stanford Cars dataset, outperforming the second-best model MattML. Comparative experiments and ablation studies further validate the effectiveness and superiority of the proposed algorithm. Codes are available at: <a href="https://github.com/204503zzw/zzw/tree/master">https://github.com/204503zzw/zzw/tree/master</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on feature disentanglement and reconstruction network based on autoencoder for few-shot fine-grained image classification

  • Ziwei Zeng,
  • Lihong Li,
  • Peixian Teng

摘要

Current few-shot fine-grained image classification models primarily rely on mining discriminative local features to enhance inter-class separability. However, these models remain insufficient in effectively addressing the challenge of substantial intra-class variation. To address the aforementioned challenges, this paper proposes a feature disentanglement and reconstruction network based on autoencoders. This network employs the disentanglement design with content and style encoders to separately process spatial information and high-level semantic information in images. It utilizes a decoder for feature reconstruction, which can effectively achieve cross-image feature alignment. By reconstructing the features, the model’s ability to represent intra-class variability in few-shot fine-grained classification tasks has been enhanced. The experimental results demonstrate that the proposed model significantly improves classification accuracy on three canonical few-shot datasets (Stanford Dogs, Stanford Cars, and CUB-200–2011). Specifically, the model achieves a 4.08% accuracy improvement in 1-shot classification tasks on the Stanford Cars dataset, outperforming the second-best model MattML. Comparative experiments and ablation studies further validate the effectiveness and superiority of the proposed algorithm. Codes are available at: https://github.com/204503zzw/zzw/tree/master.