Extracting terms from text is relevant in many applications such as document summarizing, question answering, ontology learning and many more. However, the unstructured nature of textual data poses significant hurdles that make the automatic extraction of terms from text a challenging task. Furthermore, while significant research efforts have been dedicated to term extraction from textual data, the incorporation of domain knowledge and context awareness as a way of enriching the extracted terms remains a challenge. To tackle these challenges, this study proposes an attention-based Deep Learning model for term extraction from text using the Bidirectional Encoder Representations from Transformers (BERT) language model. The model uses (1) web scraping for term extraction from web pages, (2) BERT encoder for enriching term extraction through contextualization and cosine similarity to extract domain-specific terms and (3) WordNet for adding knowledge context and disambiguation through vocabulary labeling. The proposed system was applied to five different data samples from the horticulture domain and was evaluated with various metrics including accuracy, precision, and F1-score. The experimental results indicate that the proposed model improves the domain-specific term extraction and vocabulary labeling with an accuracy of 73%, a precision of 89%, a recall of 79%, as well as F1-Score of 84%, better than the related studies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Attention-Based Deep Learning Model for Term Extraction from Text Using BERT

  • Tsitsi Zengeya,
  • Jean Vincent Fonou Dombeu,
  • Mandlenkosi Gwetu

摘要

Extracting terms from text is relevant in many applications such as document summarizing, question answering, ontology learning and many more. However, the unstructured nature of textual data poses significant hurdles that make the automatic extraction of terms from text a challenging task. Furthermore, while significant research efforts have been dedicated to term extraction from textual data, the incorporation of domain knowledge and context awareness as a way of enriching the extracted terms remains a challenge. To tackle these challenges, this study proposes an attention-based Deep Learning model for term extraction from text using the Bidirectional Encoder Representations from Transformers (BERT) language model. The model uses (1) web scraping for term extraction from web pages, (2) BERT encoder for enriching term extraction through contextualization and cosine similarity to extract domain-specific terms and (3) WordNet for adding knowledge context and disambiguation through vocabulary labeling. The proposed system was applied to five different data samples from the horticulture domain and was evaluated with various metrics including accuracy, precision, and F1-score. The experimental results indicate that the proposed model improves the domain-specific term extraction and vocabulary labeling with an accuracy of 73%, a precision of 89%, a recall of 79%, as well as F1-Score of 84%, better than the related studies.