A Self-Supervised Bearing Fault Diagnosis Method Based on Differential Transformer under Limited Label Sample Conditions
摘要
In bearing vibration signal collections, it is difficult to acquire labeled vibration data for bearings, causing challenges to the implementation of data-driven fault diagnosis. To address this issue, a mask self-supervised learning method based on a differential vision transformer (MS-DiViT) is proposed. First, raw vibration signals are converted into time-frequency images through the wavelet transform signal processing method, achieving cross-modal representations. Then, a pretext task is constructed, where the majority of time-frequency images are masked and used as input, with the original image as the reconstruction target to learn effective features. A differential vision transformer is used as the encoder, and a vision transformer is employed as the decoder of the pretraining model for unlabeled data to complete the pretext task. Finally, after the encoder training, a classification head is added, and fine-tuning is performed with labeled data to obtain the final fault diagnosis classifier. Experiments based on public datasets validate the proposed method, showing high diagnostic accuracy even when labeled data is limited.