Recently advancements in natural language processing have revolutionized dealing with massive textual data, information extraction (IE) is a natural language processing task centered around automatically retrieving organized data from unstructured or semi-structured text. Its aim is to recognize particular details, including entities, relationships and events, from textual materials to enrich structured databases or knowledge bases. Named Entity Recognition, on the other hand, is a subset of information extraction dedicated to pinpointing and classifying named entities present in text. The inability of rule-based algorithms to intelligently make decisions paved the way for research on developing algorithms that include semantic and syntactic modes of evaluation. Current work deals with Named Entity Recognition (NER) applied on real-time unstructured clinical insurance data. Using NER, keywords are extracted and classified into semantic categories. This task was extended to extract essential entities from the raw clinical text. In this paper, we propose system to integrate the real-world clinical insurance data acquired from different geographical locations as well as varied timelines. After successful Data Integration, we plan to apply NER inorder to obtain domain-specific tags, which can be used in long future. Current work presents proof of concept of dealing with the raw data acquired presently from one client. To assess the impact of various transformer models, we experimented with three transformers: BERT, DistilBERT and RoBERTa. A comparative analysis of their results in extracting pertinent information from the enormous textual data is presented in this work. Results are encouraging, and it is observed that RoBERTa outperformed the other two algorithms to deduce relevant information with a high F1 score of 90.45. Promising results achieved in this work shall tremendously pave path for fellow researchers in conducting research in the challenging area of healthcare.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Intelligent Attention-Based Transformer Models for Text Extraction: A Proof of Concept

  • Suja Sreejith Panickar,
  • Atharva Bankar,
  • Rishabh Shinde,
  • Sahil Mondal,
  • Rohit Jadhav,
  • ArunKumar Nair

摘要

Recently advancements in natural language processing have revolutionized dealing with massive textual data, information extraction (IE) is a natural language processing task centered around automatically retrieving organized data from unstructured or semi-structured text. Its aim is to recognize particular details, including entities, relationships and events, from textual materials to enrich structured databases or knowledge bases. Named Entity Recognition, on the other hand, is a subset of information extraction dedicated to pinpointing and classifying named entities present in text. The inability of rule-based algorithms to intelligently make decisions paved the way for research on developing algorithms that include semantic and syntactic modes of evaluation. Current work deals with Named Entity Recognition (NER) applied on real-time unstructured clinical insurance data. Using NER, keywords are extracted and classified into semantic categories. This task was extended to extract essential entities from the raw clinical text. In this paper, we propose system to integrate the real-world clinical insurance data acquired from different geographical locations as well as varied timelines. After successful Data Integration, we plan to apply NER inorder to obtain domain-specific tags, which can be used in long future. Current work presents proof of concept of dealing with the raw data acquired presently from one client. To assess the impact of various transformer models, we experimented with three transformers: BERT, DistilBERT and RoBERTa. A comparative analysis of their results in extracting pertinent information from the enormous textual data is presented in this work. Results are encouraging, and it is observed that RoBERTa outperformed the other two algorithms to deduce relevant information with a high F1 score of 90.45. Promising results achieved in this work shall tremendously pave path for fellow researchers in conducting research in the challenging area of healthcare.