Medical error detection remains a critical challenge for specialized language models due to limited annotated data. This paper introduces ClinAug, a lightweight text augmentation framework designed to enhance ClinicalBERT’s performance on medical error detection tasks. Using the MEDIQA-CORR 2024 MEDEC benchmark, we evaluate four Easy Data Augmentation techniques—synonym replacement, random insertion, random swap, and random deletion—adapted for clinical text. A pilot study implements synonym replacement raised Binary F1 by 0.019 and recall by 0.036, but these gains were not statistically significant (p = 0.835). These findings position ClinAug as a practical, low-cost enhancement strategy for domain-specific clinical NLP. Future work will expand ClinAug’s scope with additional techniques and larger datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ClinAug: Closing the Performance Gap in Medical Error Detection with Low-Cost Augmentation and ClinicalBERT

  • Junli Dai,
  • Yuming Li,
  • Farhaan Mirza

摘要

Medical error detection remains a critical challenge for specialized language models due to limited annotated data. This paper introduces ClinAug, a lightweight text augmentation framework designed to enhance ClinicalBERT’s performance on medical error detection tasks. Using the MEDIQA-CORR 2024 MEDEC benchmark, we evaluate four Easy Data Augmentation techniques—synonym replacement, random insertion, random swap, and random deletion—adapted for clinical text. A pilot study implements synonym replacement raised Binary F1 by 0.019 and recall by 0.036, but these gains were not statistically significant (p = 0.835). These findings position ClinAug as a practical, low-cost enhancement strategy for domain-specific clinical NLP. Future work will expand ClinAug’s scope with additional techniques and larger datasets.