Fine-Grained Fault Diagnosis in Sucker-Rod Wells via Language-Image Pretraining and Few-Shot Finetuning
摘要
Precise and fine-grained fault diagnosis in sucker-rod pumping wells is critical to ensuring production safety and operational efficiency in modern oilfields. However, existing intelligent diagnostic approaches typically rely on large-scale labeled datasets and perform poorly in distinguishing compound or severity-graded fault types under limited sample supervision. To address these challenges, we propose a vision-language diagnostic framework that integrates bootstrapped language-image pretraining with lightweight few-shot finetuning. The framework first leverages a generative pretraining paradigm based on the BLIP architecture to learn semantic alignment between dynamometer card images and structured diagnostic text. This process constructs a cross-modal embedding space through joint optimization of language modeling, image-text matching, and contrastive losses. For downstream adaptation, we apply structure-aware pruning and partial parameter freezing to reduce trainable weights to 32.76% of the original model while preserving the pretrained alignment structure. Only a small number of annotated samples are required to adapt the decoder for compound fault composition and severity-level interpretation. Experiments on real-world datasets demonstrate that the proposed method achieves over 91% macro F1 score across 26 fine-grained fault classes with as few as 150 samples per class. These results highlight the potential of language-guided multimodal models in enabling scalable, interpretable, and high-resolution diagnostics in oilfield production systems.