错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adopting Pre-trained Large Language Models for Regional Language Tasks: A Case Study

  • Harsha Gaikwad,
  • Arvind Kiwelekar,
  • Manjushree Laddha,
  • Shashank Shahare

摘要

Large language models have revolutionized the field of Natural Language Processing. While researchers have assessed their effectiveness for various English language applications, a research gap exists for their application in low-resource regional languages like Marathi. The research presented in this paper intends to fill that void by investigating the feasibility and usefulness of employing large language models for sentiment analysis in Marathi as a case study. The study gathers a diversified and labeled dataset from Twitter that includes Marathi text with opinions classified as positive, negative, or neutral. We test the appropriateness of pre-existing language models such as Multilingual BERT (M-BERT), indicBERT, and GPT-3 ADA on the obtained dataset and evaluate how they performed on the sentiment analysis task. Typical assessment metrics such as accuracy, F1 score, and loss are used to assess the effectiveness of sentiment analysis models. This research paper presents additions to the growing area of sentiment analysis in languages that have not received attention. They open up possibilities for creating sentiment analysis tools and applications specifically tailored for Marathi-speaking communities.