Journal-Level Citation Impact of Articles with Dataset Links in Abstracts Identified Using a Generative AI Ensemble
摘要
With the advancement of open science, sharing and reuse of research data have become increasingly important, yet dataset availability is often described only in the main text and not clearly visible in abstracts. This study examines whether signaling dataset availability in abstracts is associated with citation performance at the journal level. Using Web of Science metadata, we identified 6,224 articles published in 2023 across 261 journals whose abstracts contained URLs. Three generative AI systems (Copilot, ChatGPT, and Gemini) classified whether the URLs referred to research datasets using a majority-voting scheme. Human validation (n = 200) showed high precision (0.95), substantial agreement (Cohen’s \(\kappa \) = 0.75), and high accuracy (0.88). Journal-level citation differences between Data-Linked Articles (DLAs) and Non-Data-Linked Articles (NDLAs) were tested using the Mann–Whitney U test with Benjamini–Hochberg correction. Across all journals, the weighted average median difference was 4.05 citations (95% CI: 2.64–5.63) in favor of DLAs, increasing to 10.01 citations among statistically significant journals. Although causal inference is beyond the scope of this study, we do not attempt causal interpretation. The findings indicate a journal-contextual association between abstract-level dataset linking and higher citation performance, robust to unanimous agreement across all three AI systems. To support transparency and reproducibility, the analyzed abstracts, prompts, and classification outputs are publicly available at: https://github.com/opendata-study/Data-Linked-Articles .