错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Information Extraction for Design of a Multi-feature Hybrid Approach for Pronominal Anaphora Resolution in a Low Resource Language

  • Shreya Agarwal,
  • Prajna Jha,
  • Ali Abbas,
  • Tanveer J. Siddiqui

摘要

In this paper, we present a hybrid approach for anaphora resolution in Hindi Discourse which is a low resource language. We propose a rule-based module for resolving reflexive anaphoric references and a data driven approach for resolving demonstrative, relative pronouns in Hindi data. The combined module tends to resolve all three types of pronominal anaphors using syntactic features, semantic-disambiguator (named-entity and animacy), linguistic features and statistical metrics, to resolve both inter-sentential and intra-sentential anaphoric references. We evaluated our approach on Hindi tourism data and developed a corpus for developing statistical model for the same. We obtained encouraging results which holds promise and achieves accuracy of 82.9% (0.829) comparable to that of state of art approaches developed so far. The dataset used in this work contains complex, compound sentences which makes information extraction pipeline designing a challenging task, while most of the learning and non-learning approaches developed so far in Hindi have been tested and evaluated mostly on simple short stories.