A Domain-agnostic Vision Transformer Framework for Machinery Fault Diagnosis using Wavelet Time–frequency Spectra
摘要
Machinery fault diagnosis systems often exhibit degraded performance when applied across different datasets due to variations in operating conditions, machine types, and data acquisition setups. This study addresses the challenge of cross-domain generalization by developing a robust framework capable of transferring learned knowledge across heterogeneous domains without retraining.
MethodsA hybrid approach combining Continuous Wavelet Transform (CWT) and Vision Transformers (ViT) is proposed. The CWT converts raw vibration signals into time–frequency representations that preserve physics-consistent fault characteristics, while the ViT leverages global self-attention to learn structure-level invariant features. The model is trained solely on the MaFaulDa dataset and evaluated on multiple unseen datasets in a strict zero-shot setting, without any fine-tuning or domain adaptation.
ResultsThe proposed framework achieves 96.31% accuracy on the source dataset and demonstrates strong cross-domain performance, including 76% accuracy on the ComFaulDa dataset. Furthermore, a binary classification evaluation (healthy vs. faulty) conducted across six heterogeneous datasets shows consistently high performance, indicating that the learned representations effectively generalize across different domains. Performance degradation is observed primarily in scenarios with low signal-to-noise ratios and weak fault signatures rather than due to domain mismatch.
ConclusionThe results confirm that combining physics-informed time–frequency representations with attention-based learning enables robust cross-domain fault diagnosis. The proposed framework provides a scalable and practical solution for real-world industrial deployment, where labeled data from target domains are often unavailable, while also highlighting the importance of signal quality in achieving reliable generalization.