<p>Relation extraction (RE) is a fundamental task in natural language processing, and it is of crucial significance to applications such as structured search, sentiment analysis, question answering, and summarization. RE involves the identification of relations between entities, which is a challenging task for languages with low digital resources. This task becomes intricate in languages with pronoun-dropping and subject–object–verb (SOV) word order specifications, such as Persian, due to their distinctive syntactic structures. In the context of low-resource languages like Persian, this study proposes a customized model based on linguistic properties. Leveraging pronoun-dropping and subject–object–verb (SOV) word order specifications of Persian, we introduce an innovative enhancement: a novel weighted relative positional encoding integrated into the self-attention mechanism. Moreover, we enrich context representations by infusing co-occurrence information through pointwise mutual information factors. The outcomes highlight the potential of our method in relation extraction for SOV languages, shedding light on the role of syntax in enhancing NLP tasks in such linguistic contexts. We have trained and tested our model on Persian and English RE datasets. Our experiments demonstrate that our proposed model outperforms other existing Persian relation extraction (RE) models. Furthermore, our in-depth ablation study and case study show that our system can converge faster and is less prone to overfitting.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Relation extraction with enhanced self-attention model for SOV word order languages: Persian case study

  • Ebrahim Ganjalipour,
  • Amir Hossein Refahi Sheikhani,
  • Sohrab Kordrostami,
  • Ali Asghar Hosseinzadeh

摘要

Relation extraction (RE) is a fundamental task in natural language processing, and it is of crucial significance to applications such as structured search, sentiment analysis, question answering, and summarization. RE involves the identification of relations between entities, which is a challenging task for languages with low digital resources. This task becomes intricate in languages with pronoun-dropping and subject–object–verb (SOV) word order specifications, such as Persian, due to their distinctive syntactic structures. In the context of low-resource languages like Persian, this study proposes a customized model based on linguistic properties. Leveraging pronoun-dropping and subject–object–verb (SOV) word order specifications of Persian, we introduce an innovative enhancement: a novel weighted relative positional encoding integrated into the self-attention mechanism. Moreover, we enrich context representations by infusing co-occurrence information through pointwise mutual information factors. The outcomes highlight the potential of our method in relation extraction for SOV languages, shedding light on the role of syntax in enhancing NLP tasks in such linguistic contexts. We have trained and tested our model on Persian and English RE datasets. Our experiments demonstrate that our proposed model outperforms other existing Persian relation extraction (RE) models. Furthermore, our in-depth ablation study and case study show that our system can converge faster and is less prone to overfitting.