错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Synergistic Text Annotation Based on Rule-Based Expressions and DistilBERT

  • Arafet Sbei,
  • Khaoula ElBedoui,
  • Walid Barhoumi

摘要

This study introduces a novel hybrid approach to text annotation that combines rule-based regular expressions with the pretrained neural network model DistilBERT. Given limited task-specific labeled data, regular expressions are first leveraged to efficiently annotate sentences, providing a cost-effective alternative to manual labeling. The annotated dataset then serves as training data for DistilBERT, enabling the model to learn nuanced linguistic patterns and improve upon the rule-based annotations. Results demonstrate that this pretraining strategy significantly enhances performance, achieving state-of-the-art models performance, notably those reliant solely on prompt engineering, such as the biggest large language model GPT-4. This study underscores the efficacy of integrating data-driven strategies with modern pretrained models, particularly for tasks where annotated data is scarce. The proposed method presents a promising direction for building robust and adaptable sentence annotation pipelines across diverse and resource-constrained natural language processing applications. By capitalizing on both manually crafted rules and learned representations, this hybrid approach can potentially generalize better compared to relying solely on either technique alone.