Structured Knowledge Extraction for Digital Twins: Leveraging LLMs to Analyze Tweets
摘要
This paper concentrates on the extraction of pertinent information from unstructured data, specifically analyzing textual content disseminated by users on X/Twitter. The objective is to construct an exhaustive knowledge graph by discerning implicit personal data from tweets. The gleaned information serves to instantiate a digital counterpart and establish a tailored alert mechanism aimed at shielding users from threats such as social engineering or doxing. The study assesses the efficacy of fine-tuning cutting-edge open source large language models for extracting pertinent triples from tweets. Additionally, it delves into the concept of digital counterparts within the realm of cyber threats and presents relevant works in information extraction. The methodology encompasses data acquisition, relational triple extraction, large language model fine-tuning, and subsequent result evaluation. Leveraging a X/Twitter dataset, the study scrutinizes the challenges inherent in user-generated data. The outcomes underscore the precision of the extracted triples and the discernible personal traits gleaned from tweets.