<p>Establishing a knowledge graph for the fault diagnosis of turbine generator sets enables the comprehensive integration of equipment lifecycle data for operation and maintenance. The existing challenges in recognizing fault entities are addressed, including the absence of publicly available annotated corpus datasets with annotations for named entities, the heterogeneity of data from multiple sources in industrial scenarios, and the difficulty of extracting associative weight features for specialized vocabularies, making the construction of a knowledge graph difficult achieving. To address these issues, an annotated corpus dataset for named entity recognition in turbine generator set fault diagnosis based on publicly available information is constructed, and a named entity recognition method is proposed by integrating Multi-headed Self-attention (MHSA) with Robustly Optimized BERT Approach (RoBERTa), Bidirectional Long Short-Term Memory (BiLSTM), and Conditional Random Field (CRF). Specifically, the input annotated utterances undergo pre-trained in the RoBERTa layer to convert them into a sequence of word vectors, which are subsequently input into the MHSA_BiLSTM layer to extract long-range dependency information and global semantic features. The output vector from the MHSA_BiLSTM layer is fed into the CRF layer to determine the optimal global sequence. The experimental results indicate that the proposed method significantly surpasses the other named entity recognition techniques in accurately identifying fault entity categories within the specialized field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Named entity recognition method of turbine generator set fault based on MHSA and RoBERTa-BiLSTM-CRF

  • Yintao Yang,
  • Changfeng Yan,
  • Jiang Wang,
  • Jianxiong Kang,
  • Jiawei Lu

摘要

Establishing a knowledge graph for the fault diagnosis of turbine generator sets enables the comprehensive integration of equipment lifecycle data for operation and maintenance. The existing challenges in recognizing fault entities are addressed, including the absence of publicly available annotated corpus datasets with annotations for named entities, the heterogeneity of data from multiple sources in industrial scenarios, and the difficulty of extracting associative weight features for specialized vocabularies, making the construction of a knowledge graph difficult achieving. To address these issues, an annotated corpus dataset for named entity recognition in turbine generator set fault diagnosis based on publicly available information is constructed, and a named entity recognition method is proposed by integrating Multi-headed Self-attention (MHSA) with Robustly Optimized BERT Approach (RoBERTa), Bidirectional Long Short-Term Memory (BiLSTM), and Conditional Random Field (CRF). Specifically, the input annotated utterances undergo pre-trained in the RoBERTa layer to convert them into a sequence of word vectors, which are subsequently input into the MHSA_BiLSTM layer to extract long-range dependency information and global semantic features. The output vector from the MHSA_BiLSTM layer is fed into the CRF layer to determine the optimal global sequence. The experimental results indicate that the proposed method significantly surpasses the other named entity recognition techniques in accurately identifying fault entity categories within the specialized field.