错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

biblioverlap: an R package for document matching across bibliographic datasets

  • Gabriel Alves Vieira,
  • Jacqueline Leta

摘要

Bibliographic databases have long been a cornerstone of scientometrics research, and new information sources have prompted several comparative studies between them. Such studies often employ document-level matching procedures to identify overlaps in the corpus of each database and assess their coverage. However, despite being increasingly relevant in comparative studies, such a type of analysis still lacks an open-source tool to automate it. To fill this gap, we have developed an R package called biblioverlap, which implements a hybrid matching approach using a unique identifier and a selection of ubiquitous bibliographic fields to establish document co-occurrence. It supports data analysis from a broad range of secondary sources and can be used for comparing databases and assessing document overlap in virtually any bibliographic dataset, which can be insightful for various research questions. This paper presents the biblioverlap tool, details the matching procedure’s implementation, and uses an example dataset containing records from the Federal University of Rio de Janeiro to illustrate the package’s built-in functionality.