错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Generation of a Large Multiple-Choice Question-Answer Corpus

  • David Kauchak,
  • Vivien Song,
  • Prashant Mishra,
  • Gondy Leroy,
  • Phil Harber,
  • Stephen Rains,
  • John Hamre,
  • Nick Morgenstein

摘要

Large corpora with fine-grained metrics for difficulty and understandability are a critical resource for developing algorithms and tools to create more informative content. We introduce a new approach for automatically generating a large corpus of health-related content with associated multiple-choice questions using Google’s related questions and ChatGPT, including two new algorithms for generating potential wrong answers. We compare both the question quality as well as the suggested wrong answers using automated metrics and user studies. Overall, we find both algorithms generate reasonable questions that are complementary. Google questions use more accessible language and are easier to answer while ChatGPT questions appear easier, but are more difficult to answer and have better coverage over the entire text. For wrong answer generation, we find ChatGPT produces higher quality wrong answers that are more likely to be good distractors and are more closely related to the text content than our corpus-based approaches. We recommend both questions as options for studies with wrong answers generated by ChatGPT.