Large-scale and high-quality corpora are essential for evaluating question-answer generation (QAG) models, especially in low-resource languages such as Vietnamese. Within the education domain, QAG holds significant potential for practical applications; however, research in this field remains limited. To fill the gap, this paper introduces ViEduQA, a novel dataset designed to advance QAG in Vietnamese educational contexts. The dataset, covering four high school subjects (Biology, Geography, History, and Civic Education), consists of a QAG task that generates question-answer pairs directly from educational texts. We explore the capabilities of pre-trained language models (PLMs) like ViT5 and BARTPho and large language models (LLMs) such as SeaLLMs and Qwen2, leveraging fine-tuning and instruction-tuning methods. Finally, our analysis shows that the instruction tuning method has the potential to enhance the performance of QAG models, though it requires additional data. We also provide a system demonstrating how our QAG models function in education.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViEduQA: A New Vietnamese Dataset for Question Answer Generation in Education

  • Truong-Phuc Nguyen,
  • Huu-Loi Le,
  • Pham Quoc-Hung,
  • Nong Quang Huy,
  • Xuan-Hieu Phan,
  • Minh-Tien Nguyen

摘要

Large-scale and high-quality corpora are essential for evaluating question-answer generation (QAG) models, especially in low-resource languages such as Vietnamese. Within the education domain, QAG holds significant potential for practical applications; however, research in this field remains limited. To fill the gap, this paper introduces ViEduQA, a novel dataset designed to advance QAG in Vietnamese educational contexts. The dataset, covering four high school subjects (Biology, Geography, History, and Civic Education), consists of a QAG task that generates question-answer pairs directly from educational texts. We explore the capabilities of pre-trained language models (PLMs) like ViT5 and BARTPho and large language models (LLMs) such as SeaLLMs and Qwen2, leveraging fine-tuning and instruction-tuning methods. Finally, our analysis shows that the instruction tuning method has the potential to enhance the performance of QAG models, though it requires additional data. We also provide a system demonstrating how our QAG models function in education.