<p>Transformers have recently made significant achievements in image classification, mainly due to the pivotal role played by the self-attention mechanism. However, Transformers tend to have deeper and larger network structures, higher computational complexity, lack of local feature learning capability built in convolutional neural networks (ConvNets), and limitations in scalable parallelization computation. To address this issue, we incorporate ConvNets design elements into Transformers to realize complementary benefits, using shallower and wider network layouts to achieve depth-complexity decoupling and parallel adaptability. In this paper, we propose an efficient scalable multi-scale conv-attentional vision transformer network backbone for few-shot image classification, called PathViT. We redesign the internal space structure of Transformers and propose a novel Factorized Multi-head Self-attention with Enhanced-side (FMSA). We construct a fusion block based on multi-scale convolution and FMSA, which can fully exploit the advantages of local inductive biases in ConvNets and global context by the self-attention mechanism to improve feature representation and model generalization. In addition, we propose a Fused Channel-Spatial Global Attention (FGA) to enhance the feature learning capability of multi-scale convolution, and implement network design approaches that combine a hierarchical architecture with a depth of only 14 layers and multi-pathway parallelism for better flexibility and scalability. The experimental results show that PathViT achieves outstanding performance on the Oxford Flowers-102, Stanford Cars, Cifar10, Cifar100, and TinyImageNet datasets, demonstrating remarkable competitive advantages compared to other state-of-the-art (SOTA) models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PathViT: an efficient scalable multi-scale conv-attentional vision transformer for few-shot image classification

  • Zhonghua Fan,
  • Dongbai Sun,
  • Hongying Yu,
  • Weidong Zhang

摘要

Transformers have recently made significant achievements in image classification, mainly due to the pivotal role played by the self-attention mechanism. However, Transformers tend to have deeper and larger network structures, higher computational complexity, lack of local feature learning capability built in convolutional neural networks (ConvNets), and limitations in scalable parallelization computation. To address this issue, we incorporate ConvNets design elements into Transformers to realize complementary benefits, using shallower and wider network layouts to achieve depth-complexity decoupling and parallel adaptability. In this paper, we propose an efficient scalable multi-scale conv-attentional vision transformer network backbone for few-shot image classification, called PathViT. We redesign the internal space structure of Transformers and propose a novel Factorized Multi-head Self-attention with Enhanced-side (FMSA). We construct a fusion block based on multi-scale convolution and FMSA, which can fully exploit the advantages of local inductive biases in ConvNets and global context by the self-attention mechanism to improve feature representation and model generalization. In addition, we propose a Fused Channel-Spatial Global Attention (FGA) to enhance the feature learning capability of multi-scale convolution, and implement network design approaches that combine a hierarchical architecture with a depth of only 14 layers and multi-pathway parallelism for better flexibility and scalability. The experimental results show that PathViT achieves outstanding performance on the Oxford Flowers-102, Stanford Cars, Cifar10, Cifar100, and TinyImageNet datasets, demonstrating remarkable competitive advantages compared to other state-of-the-art (SOTA) models.