The growing volume of unstructured clinical text presents significant challenges in healthcare, particularly for languages with limited natural language processing (NLP) resources, such as Estonian. This study evaluates the performance of three NLP tools – Microsoft Azure Text Analytics for Health, Amazon Comprehend Medical, and MedCat – for their ability to process Estonian clinical unstructured text. The evaluation focuses on Named Entity Recognition (NER), contextual relationship extraction, assertion detection, and SNOMED CT terminology mapping. Key findings indicate that while Azure was most consistent in extracting contextual relationships, it exhibited classification errors. Amazon extracted fewer entities but tended to be more accurate with found relationships and preserving logical sentence structures by segmenting text into meaningful phrases, whereas MedCat was reliable at mapping SNOMED CT concepts but lacked assertion detection and contextual understanding. A key challenge identified was translation-based preprocessing, where inaccuracies in DeepL translations distorted clinical descriptions, affecting the integrity of context and terminology. To address terminology mismatches, integrating the SNOMED terminology server with NLP pipelines can ensure correct mapping of the Estonian Edition of SNOMED CT, improving interoperability in healthcare. Future work should prioritize training domain-specific models on Estonian clinical text to enhance accuracy and reduce reliance on translations. This study underscores the potential for NLP-driven automation in Estonian healthcare while identifying crucial areas for further development to ensure precision and reliability in clinical text processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Artificial Intelligence in the Analysis of Estonian Unstructured Clinical Text

  • Ken Kruuser,
  • Igor Bossenko,
  • Ahti Lohk,
  • Gunnar Piho,
  • Peeter Ross

摘要

The growing volume of unstructured clinical text presents significant challenges in healthcare, particularly for languages with limited natural language processing (NLP) resources, such as Estonian. This study evaluates the performance of three NLP tools – Microsoft Azure Text Analytics for Health, Amazon Comprehend Medical, and MedCat – for their ability to process Estonian clinical unstructured text. The evaluation focuses on Named Entity Recognition (NER), contextual relationship extraction, assertion detection, and SNOMED CT terminology mapping. Key findings indicate that while Azure was most consistent in extracting contextual relationships, it exhibited classification errors. Amazon extracted fewer entities but tended to be more accurate with found relationships and preserving logical sentence structures by segmenting text into meaningful phrases, whereas MedCat was reliable at mapping SNOMED CT concepts but lacked assertion detection and contextual understanding. A key challenge identified was translation-based preprocessing, where inaccuracies in DeepL translations distorted clinical descriptions, affecting the integrity of context and terminology. To address terminology mismatches, integrating the SNOMED terminology server with NLP pipelines can ensure correct mapping of the Estonian Edition of SNOMED CT, improving interoperability in healthcare. Future work should prioritize training domain-specific models on Estonian clinical text to enhance accuracy and reduce reliance on translations. This study underscores the potential for NLP-driven automation in Estonian healthcare while identifying crucial areas for further development to ensure precision and reliability in clinical text processing.