<p>Distant supervision automatically generates large-scale annotated data for relation extraction by aligning texts with knowledge bases, reducing the dependence on human annotation. However, distant supervision relation extraction inevitably introduces label noise, including false positive (FP) noise caused by neglecting sentence meanings and false negative (FN) noise due to the incompleteness of knowledge bases. Previous sentence-level methods mainly focus on the FP noise and ignore the FN noise, which induces severe misleading in both training and testing procedures. To address this issue, we propose a novel two-stage sentence-level noise reduction framework that explicitly tackles both the FN and FP problems. At stage one, we perform noise-filtering with label entailment, which filters out the FN noise before training through semantic matching between the negative instance and every relation label. At stage two, we propose robust training with collaborative denoising, which dynamically removes FP noise during training by maintaining two relation classifiers simultaneously and enabling them to learn useful knowledge from each other. Experimental results show that our method achieves significant improvements over previous state-of-the-art methods on two widely-used benchmarks. For example, a 2.55% F1 score improvements on NYT-10 dataset with BiLSTM implementation is achieved. Moreover, we validate the effectiveness of our method in reducing both the FN and FP noise.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distant supervised relation extraction with label entailment and collaborative denoising

  • Tingyu Xie,
  • Qi Li,
  • Gaoang Wang,
  • Hongwei Wang

摘要

Distant supervision automatically generates large-scale annotated data for relation extraction by aligning texts with knowledge bases, reducing the dependence on human annotation. However, distant supervision relation extraction inevitably introduces label noise, including false positive (FP) noise caused by neglecting sentence meanings and false negative (FN) noise due to the incompleteness of knowledge bases. Previous sentence-level methods mainly focus on the FP noise and ignore the FN noise, which induces severe misleading in both training and testing procedures. To address this issue, we propose a novel two-stage sentence-level noise reduction framework that explicitly tackles both the FN and FP problems. At stage one, we perform noise-filtering with label entailment, which filters out the FN noise before training through semantic matching between the negative instance and every relation label. At stage two, we propose robust training with collaborative denoising, which dynamically removes FP noise during training by maintaining two relation classifiers simultaneously and enabling them to learn useful knowledge from each other. Experimental results show that our method achieves significant improvements over previous state-of-the-art methods on two widely-used benchmarks. For example, a 2.55% F1 score improvements on NYT-10 dataset with BiLSTM implementation is achieved. Moreover, we validate the effectiveness of our method in reducing both the FN and FP noise.