The foundation of effective analysis and decision-making in Cyber Threat Intelligence (CTI) relies on a robust knowledge basis of concepts (threat actors, tools, malwares,…) and relationships between them (targets, located at,…) enabling analysts to identify, analyze, and mitigate potential threats. STIX (Structured Threat Information Expression) is a serialized structured format that aims to facilitate the exchange of cyber threat intelligence (CTI) data between organizations by interoperability between cybersecurity tools. To enable STIX objects generation from unstructured documents containing emerging objects (threats, actors,…), it is required to automate extracting concepts and the relationships between. These extractions can be tackled as supervised machine learning (ML) tasks. Supervised ML models training requires the availability of a high-quality annotated corpus with enough concepts and relationships. Manual annotation by experts is costly due to the semantic ambiguity between certain concepts and relationships and their scarcity in the documents. In this paper we present a method based on large language models (LLM) to automate the creation of the annotated corpus. The LLMs twice. First, they are used to automatically annotate a corpus from attack using a couple of prompts fitted to extract STIX concepts and relationships. As some concepts or relationships underrepresented, the corpus is completed by a second synthetic one obtained by asking LLMs to generate new documents inspired by existing one and containing concepts and relation provided in the prompt. Our evaluation of extraction models shows the effectiveness of this approach as the quality metrics are satisfactory.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM Based Data Annotation and Augmentation for NER and Relationship Extraction Models Enhancement

  • Skander Soltani,
  • Elias Limouni

摘要

The foundation of effective analysis and decision-making in Cyber Threat Intelligence (CTI) relies on a robust knowledge basis of concepts (threat actors, tools, malwares,…) and relationships between them (targets, located at,…) enabling analysts to identify, analyze, and mitigate potential threats. STIX (Structured Threat Information Expression) is a serialized structured format that aims to facilitate the exchange of cyber threat intelligence (CTI) data between organizations by interoperability between cybersecurity tools. To enable STIX objects generation from unstructured documents containing emerging objects (threats, actors,…), it is required to automate extracting concepts and the relationships between. These extractions can be tackled as supervised machine learning (ML) tasks. Supervised ML models training requires the availability of a high-quality annotated corpus with enough concepts and relationships. Manual annotation by experts is costly due to the semantic ambiguity between certain concepts and relationships and their scarcity in the documents. In this paper we present a method based on large language models (LLM) to automate the creation of the annotated corpus. The LLMs twice. First, they are used to automatically annotate a corpus from attack using a couple of prompts fitted to extract STIX concepts and relationships. As some concepts or relationships underrepresented, the corpus is completed by a second synthetic one obtained by asking LLMs to generate new documents inspired by existing one and containing concepts and relation provided in the prompt. Our evaluation of extraction models shows the effectiveness of this approach as the quality metrics are satisfactory.