Information retrieval in the healthcare industry has been improved by recent developments in Named Entity Recognition (NER); nevertheless, issues including inconsistent nomenclature, contextual variability, and the requirement for sizable, domain-specific datasets still need to be addressed. The accuracy of NER in medical texts has been improved by recent techniques, such as transformer-based models like BERT and BioBERT, although these issues continue to prevail. Furthermore, while there has been some progress in integrating deep learning techniques with conventional rule-based systems, the complexity of medical terminology frequently yields less than ideal outcomes. As BERT and deep learning techniques have advanced, the study still used Sketch Engine and BeautifulSoup to produce more specialized corpora that were suited to this specific medical keywords. This allowed to ensure that the dataset which was used for the research was more contextually relevant and directly targeted. The research investigates the linguistic patterns in drug-related medical texts by constructing two specialized corpora, using BeautifulSoup and Sketch Engine. Focusing on the medical conditions based on various sets of keywords, the corpus generated by BeautifulSoup is based on the following keyword combinations [“sinus”, “allergy”, and “migraine”]. Whereas Sketch Engine worked with three sets of keywords, namely [“sinus”, “allergy”, “migraine”], [“Bilirubin”, “Hepatitis”, “Yellowing (of skin/eyes)”], [“Lens opacity”, “Blurred vision”, “Eye surgery”]. Four corpora were generated from the NHS website—three using Sketch Engine (resulting in 28,433 words, 5974 words, 524 words, respectively) and the other through BeautifulSoup with 395,244 characters. Similarly, corpora were acquired from WebMD yielding 7904 words via BeautifulSoup, alongside a distinct dataset created using Sketch Engine. The purpose of these corpora is to facilitate the future development of Named Entity Recognition (NER), which is a crucial task in medical text mining. This paper intends to discover and categorize drug-related entities by extracting and evaluating data from reliable health information sources. This would enable more accurate and efficient information retrieval in healthcare applications. To guarantee a comprehensive corpus, the methodology combines sophisticated corpus linguistics tools with conventional online scraping techniques. As per preliminary results, combining these methods improves the recognition of important medical phrases and how they are used in context which has great promise for enhancing NER models in the field of pharmacology. Current work contributes to improved health information management and clinical decision-making processes by laying the groundwork for future research aimed at improving NER approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Creating Drugs Related Corpora for Efficient Drug Named Entity Recognition

  • Suja Sreejith Panickar,
  • Lisa Verma,
  • Riya Agrawal,
  • Vaidehi Jadhav,
  • Sakshi Gupta

摘要

Information retrieval in the healthcare industry has been improved by recent developments in Named Entity Recognition (NER); nevertheless, issues including inconsistent nomenclature, contextual variability, and the requirement for sizable, domain-specific datasets still need to be addressed. The accuracy of NER in medical texts has been improved by recent techniques, such as transformer-based models like BERT and BioBERT, although these issues continue to prevail. Furthermore, while there has been some progress in integrating deep learning techniques with conventional rule-based systems, the complexity of medical terminology frequently yields less than ideal outcomes. As BERT and deep learning techniques have advanced, the study still used Sketch Engine and BeautifulSoup to produce more specialized corpora that were suited to this specific medical keywords. This allowed to ensure that the dataset which was used for the research was more contextually relevant and directly targeted. The research investigates the linguistic patterns in drug-related medical texts by constructing two specialized corpora, using BeautifulSoup and Sketch Engine. Focusing on the medical conditions based on various sets of keywords, the corpus generated by BeautifulSoup is based on the following keyword combinations [“sinus”, “allergy”, and “migraine”]. Whereas Sketch Engine worked with three sets of keywords, namely [“sinus”, “allergy”, “migraine”], [“Bilirubin”, “Hepatitis”, “Yellowing (of skin/eyes)”], [“Lens opacity”, “Blurred vision”, “Eye surgery”]. Four corpora were generated from the NHS website—three using Sketch Engine (resulting in 28,433 words, 5974 words, 524 words, respectively) and the other through BeautifulSoup with 395,244 characters. Similarly, corpora were acquired from WebMD yielding 7904 words via BeautifulSoup, alongside a distinct dataset created using Sketch Engine. The purpose of these corpora is to facilitate the future development of Named Entity Recognition (NER), which is a crucial task in medical text mining. This paper intends to discover and categorize drug-related entities by extracting and evaluating data from reliable health information sources. This would enable more accurate and efficient information retrieval in healthcare applications. To guarantee a comprehensive corpus, the methodology combines sophisticated corpus linguistics tools with conventional online scraping techniques. As per preliminary results, combining these methods improves the recognition of important medical phrases and how they are used in context which has great promise for enhancing NER models in the field of pharmacology. Current work contributes to improved health information management and clinical decision-making processes by laying the groundwork for future research aimed at improving NER approaches.