Sanskrit is a culturally rich but low resource language. Presently, it is under represented and there is no digitized corpus available for Sanskrit-Hindi language pair publicly. We present a dataset, SHiTraD that consists of 45,500 parallel Sanskrit-Hindi sentences, collected from various sources online or offline through web or scanned from books. The source-target paired sentences were created using manual translations. This dataset is then used to train models and evaluate the translations and the results are shown using BLEU score evaluation metric comparing three architecture models 1. Encoder decoder with attention mechanism, 2. Self attention based transformer mechanism and 3. Multi-head based transformer mechanism.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SHiTraD: Sanskrit-Hindi Translation Dataset

  • Nandini Sethi,
  • Ayushi,
  • Amita Dev,
  • Poonam Bansal

摘要

Sanskrit is a culturally rich but low resource language. Presently, it is under represented and there is no digitized corpus available for Sanskrit-Hindi language pair publicly. We present a dataset, SHiTraD that consists of 45,500 parallel Sanskrit-Hindi sentences, collected from various sources online or offline through web or scanned from books. The source-target paired sentences were created using manual translations. This dataset is then used to train models and evaluate the translations and the results are shown using BLEU score evaluation metric comparing three architecture models 1. Encoder decoder with attention mechanism, 2. Self attention based transformer mechanism and 3. Multi-head based transformer mechanism.