Analyzing Biases in Popular Answer Selection Datasets on Neural-Based QA Models
摘要
The amount of information available on the internet has increased exponentially over the past decade. This digitization leads to the need of automated answering system to extract useful information from different sources. Due to the high demand of automated answering systems, many large-scale QA datasets and QA models have been introduced to the field to satisfy this need. In this work, we aim to explore and shed light upon the composition of the most popular QA datasets by comparing them through statistical distribution analyses and their biases. We collect multiple open QA datasets which cover different aspects of QA features, and highlight the differences of each QA dataset and its bias by comparing its effect on multiple baseline neural QA models. Our goal is to provide a clear understanding on the relationship of QA datasets and QA models, and offer a solid foundation for future research to enhance this growing field.