Traditional methods for Event Argument Extraction (EAE) typically require large amounts of domain-specific and high-quality annotated data from human experts, which is time-consuming and expensive. Therefore, it is crucial to fully utilize publicly available resources to improve model performance. In this paper, we propose IDEAE, an effective and data-efficient model that formulates EAE as a conditional generation task, trained using instruction tuning to transfer prior knowledge. Given a manually designed instruction and text, IDEAE learns to describe the events mentioned in the text in natural language following the given template, from which a deterministic algorithm extracts the final event arguments. Additionally, to utilize data efficiently, we introduce a novel data augmentation method combining rule-based and Chain-of-Thought (CoT) augmentation, distilling knowledge from large language models into smaller models. Extensive experiments on two benchmark datasets demonstrate the effectiveness of IDEAE, achieving state-of-the-art performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Instruction Tuning with Data Augmentation for Event Argument Extraction

  • Chengfei Wang,
  • Xiangyu Wang,
  • Huanhuan Chen

摘要

Traditional methods for Event Argument Extraction (EAE) typically require large amounts of domain-specific and high-quality annotated data from human experts, which is time-consuming and expensive. Therefore, it is crucial to fully utilize publicly available resources to improve model performance. In this paper, we propose IDEAE, an effective and data-efficient model that formulates EAE as a conditional generation task, trained using instruction tuning to transfer prior knowledge. Given a manually designed instruction and text, IDEAE learns to describe the events mentioned in the text in natural language following the given template, from which a deterministic algorithm extracts the final event arguments. Additionally, to utilize data efficiently, we introduce a novel data augmentation method combining rule-based and Chain-of-Thought (CoT) augmentation, distilling knowledge from large language models into smaller models. Extensive experiments on two benchmark datasets demonstrate the effectiveness of IDEAE, achieving state-of-the-art performance.