Large language models (LLMs) have significant potential to improve education. However, adopting these models in regions with limited resources, where mobile devices like smartphones are the only technology available oftentimes, faces technical challenges, especially in environments without connectivity. Given the lack of research exploring the capabilities of LLMs compatible with disconnected smartphones, this study evaluated the accuracy, completeness, and readability of these LLMs’ responses to educational questions in Portuguese. The Gemma, RedPajama, and Qwen models were analyzed through 18 objective questions, which encompassed the six knowledge dimensions of Bloom’s taxonomy for a comprehensive evaluation, with 196 human evaluators assessing their answers. Both the Qwen and Gemma models stood out, demonstrating high standards in correctness, while RedPajama exhibited inconsistencies in Portuguese language support. Based on human evaluators’ assessment, this research highlighted the feasibility of exploring LLMs compatible with disconnected smartphones in educational contexts with limited resources, despite also highlighting points for optimizations. Thus, this study concludes that to expand AI-enhanced education, it is necessary to develop solutions adapted to different languages and contexts, considering human evaluation to ensure high-quality responses to educational questions even in environments with no connectivity and other resource-constrained settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Large Language Model Quality in Resource-Constrained Environments: An Educational Stakeholders’ Survey on Accuracy, Completeness, and Readability in Brazil

  • Aristoteles Barros,
  • Mateus Monteiro,
  • Luiz Rodrigues,
  • Diego Dermeval,
  • Seiji Isotani,
  • Ig Ibert Bittencourt

摘要

Large language models (LLMs) have significant potential to improve education. However, adopting these models in regions with limited resources, where mobile devices like smartphones are the only technology available oftentimes, faces technical challenges, especially in environments without connectivity. Given the lack of research exploring the capabilities of LLMs compatible with disconnected smartphones, this study evaluated the accuracy, completeness, and readability of these LLMs’ responses to educational questions in Portuguese. The Gemma, RedPajama, and Qwen models were analyzed through 18 objective questions, which encompassed the six knowledge dimensions of Bloom’s taxonomy for a comprehensive evaluation, with 196 human evaluators assessing their answers. Both the Qwen and Gemma models stood out, demonstrating high standards in correctness, while RedPajama exhibited inconsistencies in Portuguese language support. Based on human evaluators’ assessment, this research highlighted the feasibility of exploring LLMs compatible with disconnected smartphones in educational contexts with limited resources, despite also highlighting points for optimizations. Thus, this study concludes that to expand AI-enhanced education, it is necessary to develop solutions adapted to different languages and contexts, considering human evaluation to ensure high-quality responses to educational questions even in environments with no connectivity and other resource-constrained settings.