One of the big achievements by the SaGaR is the introduction of the first-ever hybrid designed Awadhi part-of-speech tagger (POST) using an innovative approach. One of the unique challenges was that there were few digital resources available for Awadhi, but the authors, to their credit, tackled this head-on. Starting with an extremely meticulous creation of the Awadhi poetry corpus, 52,691 unique words were mined one-by-one to unveil the depth of the language. The present corpus is, in this spirit, enriched with a carefully selected set of 10 tags (Adjective, Adposition, Adverb, Conjunction, Noun, Number, Postposition, Preposition, Pronoun, Verb) that altogether yield a really strong set for gaining insight into Aw. This leads the authors to a remarkable realisation of accuracy, up to 83.37% in implemented SaGaR of Awadhi poetry (sourced from Ramcharitmanas). Significantly, the system puts the same proficiency in classifying Adjectives and Prepositions in the varied sections, showing its versatile nature. Although the Number tag is missing, SaGaR’s performance does not make any effect overall. The authors experimented with a set of different tests over SaGaR, like: Executing Awadhi poems on existing Hindi POST, examining Pre 1900 and Post 1900 poems in both Hindi and Awadhi, and even experimenting with diverse datasets like Stories, Gazals, and News in both languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SaGaR-A Parts-of-Speech Tagger for Low Resource Language Awadhi

  • Hema Gaikwad,
  • Jatinderkumar R. Saini

摘要

One of the big achievements by the SaGaR is the introduction of the first-ever hybrid designed Awadhi part-of-speech tagger (POST) using an innovative approach. One of the unique challenges was that there were few digital resources available for Awadhi, but the authors, to their credit, tackled this head-on. Starting with an extremely meticulous creation of the Awadhi poetry corpus, 52,691 unique words were mined one-by-one to unveil the depth of the language. The present corpus is, in this spirit, enriched with a carefully selected set of 10 tags (Adjective, Adposition, Adverb, Conjunction, Noun, Number, Postposition, Preposition, Pronoun, Verb) that altogether yield a really strong set for gaining insight into Aw. This leads the authors to a remarkable realisation of accuracy, up to 83.37% in implemented SaGaR of Awadhi poetry (sourced from Ramcharitmanas). Significantly, the system puts the same proficiency in classifying Adjectives and Prepositions in the varied sections, showing its versatile nature. Although the Number tag is missing, SaGaR’s performance does not make any effect overall. The authors experimented with a set of different tests over SaGaR, like: Executing Awadhi poems on existing Hindi POST, examining Pre 1900 and Post 1900 poems in both Hindi and Awadhi, and even experimenting with diverse datasets like Stories, Gazals, and News in both languages.