错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Effectiveness of Large Language Models in Academic Support

  • Treesha Rani Robi Das,
  • Nusrat Jahan,
  • Sajjad Waheed,
  • Partha Adhikari,
  • Mohammad Badrul Alam Miah,
  • Md. Fuyad Al Masud,
  • A. S. M. Sanwar Hosen

摘要

Artificial intelligence (AI) is transforming education by providing innovative tools for learning support, yet the effectiveness of large language models (LLMs) for younger, bilingual learners in developing regions remains underexplored. This study evaluates five freely available LLMs—ChatGPT, Gemini, DeepSeek, NotebookLM, and Grok—focusing on their ability to summarize content, generate study notes, and foster critical thinking for 10th-grade students in Bangla and English. A dataset of sixty(60) samples, drawn from six Bangla and six English textbooks, was analyzed using both qualitative methods (expert evaluations) and quantitative measures, including syntactic complexity, semantic richness, readability, and response time. Findings reveal significant performance differences. Grok produced the most accurate and contextually relevant summaries (R \(^{2}\) = 57.22%, expert rating = 4.25/5), though with relatively slower response times. Gemini demonstrated the fastest responses (16.23 s) with moderate performance (R \(^{2}\) = 52.75%). ChatGPT achieved balanced outcomes (R \(^{2}\) = 45.64%, rating = 4.08), while NotebookLM and DeepSeek were less effective for this demographic and use case. The results offer practical guidance for educators, students, and policymakers in choosing AI tools for academic support. Specifically, Gemini is better suited for quick summarization tasks, while Grok is more effective for deeper comprehension and study acceleration. Beyond tool selection, this research contributes a methodological framework for assessing age-appropriate AI in multilingual educational contexts and promotes digital literacy in developing countries.