Dataset search has become a relevant retrieval task which, with the current state of the art, remains difficult to be evaluated. The reason for this is probably not a lack of retrieval models, but a lack of suitable test collections that combine data-related information needs, such as descriptions of inductive learning tasks, with answers, namely suitable data sets for estimating model function parameters for the learning task. We have identified Reddit as a promising source for the creation of a dataset retrieval benchmark and describe its construction in detail here. The benchmark consists of a collection of datasets in the form of 9935 metadata descriptions from Papers with Code ( https://paperswithcode.com/ ), a set of 814 topics in the form of queries or descriptions of the required dataset properties, and gold standard relevance assessments derived from Reddit discussions about which dataset is best suited. In a pilot study, we evaluate first baselines on answering test requests against the collection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Test Collection for Dataset Retrieval

  • Nikolay Kolyada,
  • Martin Potthast,
  • Benno Stein

摘要

Dataset search has become a relevant retrieval task which, with the current state of the art, remains difficult to be evaluated. The reason for this is probably not a lack of retrieval models, but a lack of suitable test collections that combine data-related information needs, such as descriptions of inductive learning tasks, with answers, namely suitable data sets for estimating model function parameters for the learning task. We have identified Reddit as a promising source for the creation of a dataset retrieval benchmark and describe its construction in detail here. The benchmark consists of a collection of datasets in the form of 9935 metadata descriptions from Papers with Code ( https://paperswithcode.com/ ), a set of 814 topics in the form of queries or descriptions of the required dataset properties, and gold standard relevance assessments derived from Reddit discussions about which dataset is best suited. In a pilot study, we evaluate first baselines on answering test requests against the collection.