Open science is being promoted globally, and researchers are encouraged to publish and share research artifacts, e.g., datasets, tools, software, code, and scholarly papers. To accelerate open science, the accessibility of research artifacts is also important. Previous studies focused on the relevance between scholarly papers to improve their accessibility; however, the relevance between research artifacts other than scholarly papers has not been investigated extensively. In this study, we focus on datasets, and verify the feasibility of estimating the relevance between datasets. Access using the relevance between datasets may be useful when it is difficult to verbalize the characteristics of the desired dataset. We implemented a method to estimate the relevance between datasets that utilizes dataset metadata expressed as text, and the relevance is estimated by using a BERT-based model. Experimental results demonstrated the feasibility of estimating the relevance between datasets. Finally, we discuss the effect of utilizing metadata for each metadata field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Estimation of Relevance Between Datasets for Enhancing Accessibility of Research Artifacts

  • Koichiro Ito,
  • Shigeki Matsubara

摘要

Open science is being promoted globally, and researchers are encouraged to publish and share research artifacts, e.g., datasets, tools, software, code, and scholarly papers. To accelerate open science, the accessibility of research artifacts is also important. Previous studies focused on the relevance between scholarly papers to improve their accessibility; however, the relevance between research artifacts other than scholarly papers has not been investigated extensively. In this study, we focus on datasets, and verify the feasibility of estimating the relevance between datasets. Access using the relevance between datasets may be useful when it is difficult to verbalize the characteristics of the desired dataset. We implemented a method to estimate the relevance between datasets that utilizes dataset metadata expressed as text, and the relevance is estimated by using a BERT-based model. Experimental results demonstrated the feasibility of estimating the relevance between datasets. Finally, we discuss the effect of utilizing metadata for each metadata field.