Abstract <p>We present a combined pipeline for knowledge-graph construction and ontology expansion. This approach creates a BIO-tagged corpus via fully automatic LLM-based pseudoannotation and introduces dedicated UNK reserve categories to capture previously unseen classes and relations. A specialized NER/RE model is trained on a 3-million-token dataset with 92 labels. This model exhibits a conservative quality profile—high precision with moderate recall—suited for safe graph enrichment: integrating the extracted facts expands the graph to ~0.98 million triples, while the expansion ratio (total inferred facts to explicit triples) increases from 2.65 to 3.52, with logical consistency preserved. UNK label pools are converted into stable synsets, enabling semiautomatic ontology expansion; 12 new classes derived from unstructured texts were added. We also demonstrate practical value for querying and analytics using an LLM + SPARQL setup.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic and Semiautomatic Methods for Domain Knowledge-Graph Construction and Ontology Expansion

  • A. P. Khalov,
  • O. M. Ataeva

摘要

Abstract

We present a combined pipeline for knowledge-graph construction and ontology expansion. This approach creates a BIO-tagged corpus via fully automatic LLM-based pseudoannotation and introduces dedicated UNK reserve categories to capture previously unseen classes and relations. A specialized NER/RE model is trained on a 3-million-token dataset with 92 labels. This model exhibits a conservative quality profile—high precision with moderate recall—suited for safe graph enrichment: integrating the extracted facts expands the graph to ~0.98 million triples, while the expansion ratio (total inferred facts to explicit triples) increases from 2.65 to 3.52, with logical consistency preserved. UNK label pools are converted into stable synsets, enabling semiautomatic ontology expansion; 12 new classes derived from unstructured texts were added. We also demonstrate practical value for querying and analytics using an LLM + SPARQL setup.