Local Feature Expansion ViT Model for Bearing Fault Diagnosis Under Noise Environment
摘要
Vision Transformer (ViT) shows a great potential in the field of bearing fault diagnosis by virtue of its multi-head self-attention mechanism. However, ViT confines the one-layered convolutional network to the feature map preparation and its accuracy is hindered by the insufficient local features caused by noise interference in practical scenarios. To address this problem, a local feature expansion based ViT, i.e., LFE-ViT, is proposed. A hybrid convolutional residual network is introduced to the embedding module to expand the local information of the bearing faults. Then, by combination of the local feature expansion network and the following multi-head self-attention mechanism, the completely local and global feature representation is achieved for the bearing fault classification. Finally, the derived ViT model is validated on the Case Western Reserve University bearing dataset. The experimental results have shown that it gives better diagnostic performance compared with existing methods under the noise environment.