The question-answering (QA) system is a crucial task in the field of natural language processing. Although there are many well-established datasets in the English domain, collecting and constructing non-English QA datasets, particularly for Chinese, remains a significant challenge due to high costs and difficulties in data acquisition. In educational technology, automated scoring systems are essential for improving both teaching efficiency and quality. However, research on QA scoring for Chinese learners is still limited, especially in terms of constructing high-quality scoring datasets. This study aims to build a QA scoring dataset specifically for Chinese learners to support research on automated Chinese answer evaluation. We collected 13,650 responses from Chinese learners, covering 60 questions related to daily life and work. To ensure annotation accuracy, manual scoring was employed, and the inter-rater reliability among five annotators was assessed using correlation coefficients, demonstrating high consistency and reliability. Additionally, we introduced the English Mohler dataset as a comparison and used four classic deep learning models as baselines. The public LXMERT model was also used to verify the effectiveness and usability of our constructed dataset. Through this dataset construction, we hope to provide a foundation for the development of automated scoring technology for Chinese, as well as support downstream tasks such as intelligent tutoring systems and personalized learning path planning, ultimately enhancing the learning experience and outcomes for Chinese learners.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dataset Construction for Learner-Based Chinese Answering Questions Scoring

  • Yaru Jiang,
  • Aishan Wumaier

摘要

The question-answering (QA) system is a crucial task in the field of natural language processing. Although there are many well-established datasets in the English domain, collecting and constructing non-English QA datasets, particularly for Chinese, remains a significant challenge due to high costs and difficulties in data acquisition. In educational technology, automated scoring systems are essential for improving both teaching efficiency and quality. However, research on QA scoring for Chinese learners is still limited, especially in terms of constructing high-quality scoring datasets. This study aims to build a QA scoring dataset specifically for Chinese learners to support research on automated Chinese answer evaluation. We collected 13,650 responses from Chinese learners, covering 60 questions related to daily life and work. To ensure annotation accuracy, manual scoring was employed, and the inter-rater reliability among five annotators was assessed using correlation coefficients, demonstrating high consistency and reliability. Additionally, we introduced the English Mohler dataset as a comparison and used four classic deep learning models as baselines. The public LXMERT model was also used to verify the effectiveness and usability of our constructed dataset. Through this dataset construction, we hope to provide a foundation for the development of automated scoring technology for Chinese, as well as support downstream tasks such as intelligent tutoring systems and personalized learning path planning, ultimately enhancing the learning experience and outcomes for Chinese learners.