错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Construction of Knowledge Graph for Personal Sensitive Data

  • Pei Li,
  • Xuejun Bai,
  • Jingyi Li,
  • Yancheng Dong,
  • Jian Yang

摘要

This article demonstrates an automated method for producing a knowledge graph for sensitive personal information by utilizing natural language processing techniques. It employs a combination of rule-based and machine learning-based methods to extract entities and relationships from textual data and represent them as a knowledge graph. The study presents two major contributions. Firstly, it achieves unsupervised extraction of entities and relationships by using pre-trained linguistic models to extract high-quality labeled embeddings. This enables efficient clustering and scoring of potential entities and relationships, even without labeled data. To enhance the accuracy and generalization ability of extracting sensitive personal information relationships, a more comprehensive set of relationship types is considered. The experimental results demonstrate that this method significantly improves entity recognition accuracy and coverage compared to traditional named entity recognition methods. Secondly, a context-aware graph fusion mechanism is introduced to merge subgraphs extracted from multiple sentences into a unified knowledge graph. This mechanism preserves important semantic information, reduces noise and inconsistency, and results in a more accurate and robust knowledge graph with improved accuracy and coverage in identifying sensitive personal information entities. According to the experimental findings, this approach leads to a significant increase in the accuracy of relationship extraction and its ability to generalize, comparable to conventional methods. These results imply that the proposed approach can effectively generate knowledge graphs for sensitive personal information with high accuracy and generalization ability, making it suitable for fields such as personal information protection and privacy security.