A vision transformer framework for power equipment hazard identification based on conditional diffusion modality enhancement and adversarial domain adaptation
摘要
The power system must maintain a long-term stable operation state, as the operating condition of power equipment is directly related to power supply reliability and public safety. This paper proposes a conditional diffusion-based modality enhancement vision transformer (GEM-ViT) combined with adversarial domain adaptation for hidden danger identification. Although multimodal large language models (MLLMs) exhibit excellent performance on ideal datasets, their generalization ability in real-world complex scenes remains challenging. This study adopts domain adaptation technology to enhance the model’s capability in extracting and aligning key features of working conditions. By improving the vision transformer encoder to capture image semantic information and enhancing the attention mechanism, the adversarial training model implicitly reduces the distribution discrepancy between the source and target domains at the feature level, thereby promoting the learning of domain-invariant feature representations. Experimental results show that the proposed model achieves 7.3% higher average task accuracy than the baseline model. Furthermore, the feature extraction index and noise robustness index are 15.6 and 22.1% higher, respectively, compared to the baseline. From a quantitative perspective, the model’s feature fusion quality and anti-interference capability are confirmed, enabling more accurate and coherent multimodal content generation. This research provides an effective technical path for improving the practicality and reliability of MLLMs in open environments.