Large language models (LLMs) have been used to generate standard assessment items, but their capacity to generate complex, innovative items remains underexplored. A case in point is collaborative problem-solving (CPS) tasks which require communication, interdependence, and knowledge co-construction. Earlier work has shown that not only are LLMs deficient in generating CPS items, but they are also not reliably accurate at evaluating examples for basic qualities of interdependence. There is also a lack of transparency in the reasoning process of LLMs due to their “black box” nature. This study tries to address these shortcomings by using a multi-agent workflow with rubric generation as an intermediate step while evaluating CPS math items. Rubrics are a plain-language form of knowledge representation. By generating rubrics, LLMs “show” their reasoning about candidate items, making their logic transparent and human-interpretable. Using the CrewAI (multi-agent) framework, three specialized LLM agents were instructed to generate, refine, and finally apply evaluation rubrics to assess three types of CPS items ( \(N=63\) ) that were known to be “good” or “bad” based on their structured interdependence. Quantitative measures of classification accuracy were compared between the multi-agent system and single-agent models that did not produce intermediate rubrics. Furthermore, the rubrics were qualitatively analyzed to determine (a) the degree to which the LLM rubrics mimicked expert thinking about CPS and (b) reflected different criteria for different types of items. Overall accuracy using the multi-agent system was not improved relative to single-agent performance and varied considerably between item types. Qualitative analysis revealed that while LLMs consistently emphasized criteria such as information dependency and communication, they struggled to distinguish task-specific nuances.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using Generated Rubrics to Provide a Window Into Item Evaluation with Multi-agent LLMs

  • Yu Wang,
  • Madhumitha Gopalakrishnan,
  • Yoav Bergner

摘要

Large language models (LLMs) have been used to generate standard assessment items, but their capacity to generate complex, innovative items remains underexplored. A case in point is collaborative problem-solving (CPS) tasks which require communication, interdependence, and knowledge co-construction. Earlier work has shown that not only are LLMs deficient in generating CPS items, but they are also not reliably accurate at evaluating examples for basic qualities of interdependence. There is also a lack of transparency in the reasoning process of LLMs due to their “black box” nature. This study tries to address these shortcomings by using a multi-agent workflow with rubric generation as an intermediate step while evaluating CPS math items. Rubrics are a plain-language form of knowledge representation. By generating rubrics, LLMs “show” their reasoning about candidate items, making their logic transparent and human-interpretable. Using the CrewAI (multi-agent) framework, three specialized LLM agents were instructed to generate, refine, and finally apply evaluation rubrics to assess three types of CPS items ( \(N=63\) ) that were known to be “good” or “bad” based on their structured interdependence. Quantitative measures of classification accuracy were compared between the multi-agent system and single-agent models that did not produce intermediate rubrics. Furthermore, the rubrics were qualitatively analyzed to determine (a) the degree to which the LLM rubrics mimicked expert thinking about CPS and (b) reflected different criteria for different types of items. Overall accuracy using the multi-agent system was not improved relative to single-agent performance and varied considerably between item types. Qualitative analysis revealed that while LLMs consistently emphasized criteria such as information dependency and communication, they struggled to distinguish task-specific nuances.