错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DTP: An Open-Domain Text Relation Extraction Method

  • Yuanzhe Qiu,
  • Chenkai Hu,
  • Kai Gao

摘要

Open-domain text relation extraction (OpenTRE) is a subfield of information extraction that focuses on extracting relational facts from open-domain corpora. Recent OpenTRE researches have shown that clustering unlabeled instances leveraging knowledge from labeled data is effective, but most of them are based on the assumption that the testing set only contains open relations, which is inconsistent with real-world scenarios where known and open relations are mixed. Therefore, a novel OpenTRE method based on Dynamic Thresholds and Pair-based Self-weighting Loss (DTP) is proposed. It performs text data processing by categorizing instances and predicting unknown relations, which can handle a more diverse range of data. Specifically, we break down OpenTRE into two stages: detecting and discovering, which makes the OpenTRE process more understandable. Wherein, sample-based dynamic threshold strategy is employed to clarify the relation boundaries. Additionally, pair-based self-weighted loss improves the capture of semantic knowledge in labeled data and guides clustering. Experimental results indicate that this method outperforms strong baseline models on two datasets and has significant improvements.