错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Domain Abbreviation Disambiguation on Vietnamese Clinical Texts in Online Processing

  • Chau Vo,
  • Hua Phung Nguyen

摘要

Readability of clinical texts in electronic medical records (EMRs) is more significant when EMRs are shared and required to be understood by both human users and computer programs. This feature is seldom reached successfully in the real world because of noises in clinical texts such as spelling errors, abbreviations, synonyms, and sentence incompleteness. Among noises, abbreviations are ubiquitous in many various short forms for many different long forms. Therefore, abbreviation disambiguation on clinical texts has been well researched worldwide for a long time. Many languages like English, German, Korean, Swedish, etc. have been considered. Recently Vietnamese clinical text analytics has been emerging due to the more popularity of EMRs. However, few works have been dedicated to abbreviation disambiguation on Vietnamese clinical texts. As one of the first works for abbreviation disambiguation on Vietnamese clinical texts, our work aims at an effective novel solution. Different from the existing works, our solution supports a cross-domain context where data shortage and imbalance exist simultaneously in online processing. It defines Nonparametric Self-Training, a parameter-free semisupervised learning algorithm, to get rid of the mismatch between the source and target domains. It also enhances the labeled dataset of one domain by adding more unlabeled data of another domain in favour of data imbalance. As a result, the proposed solution can tackle the task challenges well and outperform the others with both Accuracy and AUC of higher 90% in most experiments on the real Vietnamese clinical texts of 4 different note types in 2 different hospitals. Cleaned clinical texts can be further processed better for more readability and sharability.