错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MbAbI: A Benchmark Dataset for Malayalam Text Understanding and Reasoning

  • K. Reji Rahmath,
  • P. C. Reghu Raj,
  • P. C. Rafeeque

摘要

This paper proposes a Malayalam Dataset containing 40000 instances, intended to assist the researchers working on Question Answering systems. This data set is the first of its kind. The Facebook bAbI dataset was translated to create this Malayalam bAbI (or, MbAbI) dataset. It comprises 20 different tasks of varying complexity. The tasks include simple supporting facts to complex path finding tasks. We have machine translated these tasks from English to Malayalam and tested the baseline QA models on these datasets. We have obtained state-of-the-art results with MbAbI dataset using deep learning and transformer models. In order to benefit the low-resource Malayalam research community we have made this dataset publicly available.