错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Pre-trained Language Models for Vocabulary Alignment in the UMLS

  • Xubing Hao,
  • Rashmie Abeysinghe,
  • Jay Shi,
  • Licong Cui

摘要

The Unified Medical Language System (UMLS) Metathesaurus integrates and aligns terms from hundreds of biomedical vocabularies. In this paper, we investigate the efficacy of Pre-trained Language Models (PLMs) for vocabulary alignment in the UMLS Metathesaurus. We frame the problem as two Natural Language Processing tasks: Text Classification and Text Generation. We fine-tune four opensource cutting-edge PLMs including BERT and RoBERTa, GPT-2, and BLOOM. Experiments show that the best model is RoBERTa achieving a precision, recall, and F1 score of 0.965, 0.940, and 0.952 respectively. In addition, incorporation of contextual information in the inputs improves the model performance in the Text Classification task, albeit with a limited impact on the Text Generation task. Domain expert evaluation of 100 randomly selected instances generated by the best model revealed that 78 of them as valid synonymous terms, indicating the promise of PLMs in enhancing the mapping quality of the UMLS Metathesaurus.