<p>The development of explanations for scientific phenomena is crucial in science assessment. However, the scoring of students’ written explanations is a challenging and resource-intensive process. Large language models (LLMs) have demonstrated the potential to address these challenges, particularly when the explanations are written in English, an alphabetic language. It remains unknown whether this approach can be applied to other logographic languages. This study thus explores the potential of fine-tuning ChatGPT, one advanced LLM, to automatically score scientific explanations written in Chinese. We collected and automatically scored student responses to seven scientific explanation tasks in Chinese, and examined the relationship between scoring accuracy and reasoning complexity with Kendall correlation. Finally, a qualitative analysis was conducted to explore how linguistic features influence scoring accuracy. The results indicate that through domain-specific adaptation, the fine-tuned ChatGPT can accurately score students’ written explanations in Chinese. However, scoring accuracy correlates with reasoning complexity, showing a negative correlation for lower-level responses and a positive one for higher-level responses. The model tends to overrate complex reasoning for low-level responses with complex sentence structures and underrate high-level responses, using generalizing, summarizing, or simple causal reasoning. These opposing correlations are associated with different linguistic features. The comprehensiveness of student responses is often in tension with the simplicity and clarity of language structure in terms of scoring accuracy. For lower-level responses, simplicity and clarity are prioritized, leading to more accurate scores for simpler and shorter responses. For higher-level responses, comprehensiveness is prioritized, resulting in more accurate scores for long and information-rich responses. These findings demonstrate the effectiveness of LLMs in automatic scoring within a Chinese context and highlight the importance of considering linguistic features and reasoning complexity in developing and fine-tuning automatic scoring models for educational assessments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese

  • Jie Yang,
  • Ehsan Latif,
  • Yuze He,
  • Xiaoming Zhai

摘要

The development of explanations for scientific phenomena is crucial in science assessment. However, the scoring of students’ written explanations is a challenging and resource-intensive process. Large language models (LLMs) have demonstrated the potential to address these challenges, particularly when the explanations are written in English, an alphabetic language. It remains unknown whether this approach can be applied to other logographic languages. This study thus explores the potential of fine-tuning ChatGPT, one advanced LLM, to automatically score scientific explanations written in Chinese. We collected and automatically scored student responses to seven scientific explanation tasks in Chinese, and examined the relationship between scoring accuracy and reasoning complexity with Kendall correlation. Finally, a qualitative analysis was conducted to explore how linguistic features influence scoring accuracy. The results indicate that through domain-specific adaptation, the fine-tuned ChatGPT can accurately score students’ written explanations in Chinese. However, scoring accuracy correlates with reasoning complexity, showing a negative correlation for lower-level responses and a positive one for higher-level responses. The model tends to overrate complex reasoning for low-level responses with complex sentence structures and underrate high-level responses, using generalizing, summarizing, or simple causal reasoning. These opposing correlations are associated with different linguistic features. The comprehensiveness of student responses is often in tension with the simplicity and clarity of language structure in terms of scoring accuracy. For lower-level responses, simplicity and clarity are prioritized, leading to more accurate scores for simpler and shorter responses. For higher-level responses, comprehensiveness is prioritized, resulting in more accurate scores for long and information-rich responses. These findings demonstrate the effectiveness of LLMs in automatic scoring within a Chinese context and highlight the importance of considering linguistic features and reasoning complexity in developing and fine-tuning automatic scoring models for educational assessments.