Scaling Language Boundaries: A Comparative Analysis of Multilingual Question-Answering Capabilities in Large Language Models
摘要
Large Language Models (LLMs), which are endowed with the capacity to understand and produce text similar to human language, have emerged as transformative entities within the field of natural language processing (NLP). Particularly, multilingual LLMs show enormous potential for bridging linguistic divides and expanding technical access to a variety of linguistic communities. In light of this, our research compares and contrasts GPT and BLOOM, two well- known LLMs, across the range of regional Indian languages. We carefully evaluate the models’ ability to produce language-specific replies with a focus on English, Hindi, and Punjabi prompts. To fully comprehend the strengths and weak- nesses of the models, this study includes measures for linguistic alignment, ac- curacy assessment, and human review of model-generated results. The accuracy analysis revealed that while BLOOM displayed a range of performance levels, GPT consistently demonstrated competency in all of the languages tested, suggesting its versatility across a wide range of languages. This comparative study has consequences that go beyond the field of AI studies. We enable informed decision-making when deploying these models for multilingual communication, content creation, and cross-cultural interactions by highlighting the strengths and shortcomings of these models. In the end, by bridging the gap between robots and other linguistic communities, this work advances both technological aspirations and societal inclusion.