Beyond existence checks: a cross-database benchmark of bibliographic metadata consistency across academic domains and LLM-generated citations
摘要
Citation verification tools increasingly rely on scholarly databases to confirm reference authenticity, yet all existing systems employ a binary approach: checking whether a citation exists in any single source. This study introduces a cross-database benchmark for bibliographic metadata consistency, evaluating whether databases agree on the metadata values associated with verified citations. A corpus of 1,246 citations—comprising 491 AI-generated references from five large language models (Claude Opus 4.6, GPT-4.1-nano, Gemini 3 Flash, Llama 3.3 70B, and Llama 3.1 8B) and 755 database-sampled references—was systematically queried against three scholarly databases (CrossRef, OpenAlex, and Semantic Scholar) across seven metadata fields and five academic domains. Results reveal that 66.7% of all citations exhibit at least one metadata disagreement between databases, with author names showing the lowest inter-database agreement (29–39%) and DOIs the highest (100%). Manual classification of 200 disagreements indicates that 38.5% represent only legitimate metadata variants (e.g., online-first vs. print year, abbreviated vs. full author names) while 61.5% contain at least one substantive inconsistency (e.g., wrong year, missing co-authors); a publisher-page calibration check on 30 consensus cases found 27 of 29 verifiable values correct, indicating that cross-database consensus is a strong but imperfect proxy for correctness. Domain-level analysis indicates that biomedicine and social science exhibit the highest proportion of citations with at least one disagreement (75% and 72%), while CS/ML shows the lowest (61%); the domain effect is statistically significant but modest. Cross-database comparison detects metadata disagreements at a field-level rate of 21.6%—entirely invisible to single-source verification. Semantic Scholar coverage is reported after an authenticated re-collection that corrects an initial rate-limit undercount. These findings demonstrate that citation verification must move beyond existence checks toward cross-database metadata comparison, with implications for verification tool design, bibliometric research, and AI-assisted scholarly workflows.