<p>The power system must maintain a long-term stable operation state, as the operating condition of power equipment is directly related to power supply reliability and public safety. This paper proposes a conditional diffusion-based modality enhancement vision transformer (GEM-ViT) combined with adversarial domain adaptation for hidden danger identification. Although multimodal large language models (MLLMs) exhibit excellent performance on ideal datasets, their generalization ability in real-world complex scenes remains challenging. This study adopts domain adaptation technology to enhance the model’s capability in extracting and aligning key features of working conditions. By improving the vision transformer encoder to capture image semantic information and enhancing the attention mechanism, the adversarial training model implicitly reduces the distribution discrepancy between the source and target domains at the feature level, thereby promoting the learning of domain-invariant feature representations. Experimental results show that the proposed model achieves 7.3% higher average task accuracy than the baseline model. Furthermore, the feature extraction index and noise robustness index are 15.6 and 22.1% higher, respectively, compared to the baseline. From a quantitative perspective, the model’s feature fusion quality and anti-interference capability are confirmed, enabling more accurate and coherent multimodal content generation. This research provides an effective technical path for improving the practicality and reliability of MLLMs in open environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A vision transformer framework for power equipment hazard identification based on conditional diffusion modality enhancement and adversarial domain adaptation

  • Ruchao Liao,
  • Duanjiao Li,
  • Gao Liu,
  • Xiuzhen Ding,
  • Zhuojun Xie

摘要

The power system must maintain a long-term stable operation state, as the operating condition of power equipment is directly related to power supply reliability and public safety. This paper proposes a conditional diffusion-based modality enhancement vision transformer (GEM-ViT) combined with adversarial domain adaptation for hidden danger identification. Although multimodal large language models (MLLMs) exhibit excellent performance on ideal datasets, their generalization ability in real-world complex scenes remains challenging. This study adopts domain adaptation technology to enhance the model’s capability in extracting and aligning key features of working conditions. By improving the vision transformer encoder to capture image semantic information and enhancing the attention mechanism, the adversarial training model implicitly reduces the distribution discrepancy between the source and target domains at the feature level, thereby promoting the learning of domain-invariant feature representations. Experimental results show that the proposed model achieves 7.3% higher average task accuracy than the baseline model. Furthermore, the feature extraction index and noise robustness index are 15.6 and 22.1% higher, respectively, compared to the baseline. From a quantitative perspective, the model’s feature fusion quality and anti-interference capability are confirmed, enabling more accurate and coherent multimodal content generation. This research provides an effective technical path for improving the practicality and reliability of MLLMs in open environments.