Identification of the Appropriate Value to Replace the Missing Value in Incomplete Triplet Extraction from Domain Text for Domain Knowledge
摘要
A domain ontology focuses more on covering the details, information, and knowledge about a specific domain chosen. It is used to assist and ease the process of gathering information related to the chosen domain. An ontology is presented by using ontology components which consist of concepts and relations. Ontology components extraction can be based on sentence triplet extraction, where the concepts are the subject and object, and relation is the predicate. Many existing works have proposed to extract the concepts and relations from domain texts. However, most of these works are only able to extract concepts (subject and object of the sentence) and relations (terms, either verb or predicate, that relate the subject and object) presented in a single complete sentence and neglect all the incomplete sentences, where either a subject or object is absent from the sentences. This situation may cause the domain ontology is not properly presented, thus making the knowledge presented about the domain is lacking since the neglected sentences might contain some valuable information that is crucial to the domain knowledge. Manual extraction by domain experts can be used to solve the extraction of ontology components from incomplete sentences, but this method is costly, time-consuming, and prone to error. In this paper, a method to solve the issue of extracting the ontology components from incomplete sentences in domain texts is presented. The algorithm developed suggests the most suitable value that can be used to replace the missing concepts, thus improving and increasing the number of domain knowledge extraction. The efficiency and functionality of this algorithm are tested by using data sets related to the cybersecurity, technology, and health domains. The outcome of the proposed method shows a positive significance in improving the domain knowledge extraction.