Outcome-Based Education (OBE) emphasizes achieving measurable student outcomes and relies heavily on Continuous Quality Improvement (CQI) to refine teaching-learning practices. However, manual analysis of assessment data and instructor feedback for CQI consumes valuable time and effort from faculty members. To mitigate this issue, this paper evaluates the potential of Large Language Models (LLMs) to automate key CQI tasks, including summarizing multi-section instructor recommendations and generating observation reports from Course Outcome (CO) assessment data. We tested six LLMs – including GPT-3.5, GPT-4, Llama-3.2 variants, and T5 models – using metrics such as ROUGE-1, BERTScore, and SBERT Cosine Similarity. We found that GPT-3.5 consistently outperformed others in both summarization and observation generation tasks in terms of both BERTScore F1 and SBERT Cosine Similarity metrics while exhibiting at least 2.4 times faster execution time than the second best model (GPT-4). To the best of our knowledge, this is the first work reported in the literature on evaluating LLM tools for CQI in outcome-based education.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Evaluation of LLM Tools for Continuous Quality Improvement in Outcome-Based Education

  • Md Shafqatur Rahman,
  • Mahdi Redwan,
  • Ahsanur Rahman,
  • Riasat Khan,
  • Rajesh Palit

摘要

Outcome-Based Education (OBE) emphasizes achieving measurable student outcomes and relies heavily on Continuous Quality Improvement (CQI) to refine teaching-learning practices. However, manual analysis of assessment data and instructor feedback for CQI consumes valuable time and effort from faculty members. To mitigate this issue, this paper evaluates the potential of Large Language Models (LLMs) to automate key CQI tasks, including summarizing multi-section instructor recommendations and generating observation reports from Course Outcome (CO) assessment data. We tested six LLMs – including GPT-3.5, GPT-4, Llama-3.2 variants, and T5 models – using metrics such as ROUGE-1, BERTScore, and SBERT Cosine Similarity. We found that GPT-3.5 consistently outperformed others in both summarization and observation generation tasks in terms of both BERTScore F1 and SBERT Cosine Similarity metrics while exhibiting at least 2.4 times faster execution time than the second best model (GPT-4). To the best of our knowledge, this is the first work reported in the literature on evaluating LLM tools for CQI in outcome-based education.