Automatic Generation of a Large Multiple-Choice Question-Answer Corpus
摘要
Large corpora with fine-grained metrics for difficulty and understandability are a critical resource for developing algorithms and tools to create more informative content. We introduce a new approach for automatically generating a large corpus of health-related content with associated multiple-choice questions using Google’s related questions and ChatGPT, including two new algorithms for generating potential wrong answers. We compare both the question quality as well as the suggested wrong answers using automated metrics and user studies. Overall, we find both algorithms generate reasonable questions that are complementary. Google questions use more accessible language and are easier to answer while ChatGPT questions appear easier, but are more difficult to answer and have better coverage over the entire text. For wrong answer generation, we find ChatGPT produces higher quality wrong answers that are more likely to be good distractors and are more closely related to the text content than our corpus-based approaches. We recommend both questions as options for studies with wrong answers generated by ChatGPT.