Evaluating the Reporting Quality of 21,041 Randomized Controlled Trial Articles with Large Language Models: A Large-Scale Transparency Analysis
摘要
Incomplete reporting of a study’s methods and results hinders efforts to evaluate and reproduce research findings in randomized controlled trials (RCTs), leading to potential harm. CONSORT, a set of widely endorsed RCT reporting guidelines, was designed to mitigate such harm by ensuring transparency, reproducibility and safety in RCTs. Evaluating adherence to CONSORT requires a level of language understanding that has previously precluded broad systematic assessment. However, recent advances in large language models (LLMs) now make it possible to evaluate the quality of RCTs at scale. We demonstrate GPT-4o-mini, used out of the box to achieve state-of-the-art performance in evaluating RCT quality (F1 score: 0.85; precision: 0.96), with results validated by expert human annotators showing 92.24% agreement across 50 papers. Applying this tool to 21,041 open-access RCTs (1966–2024), we reveal temporal and domain trends: while overall adherence to CONSORT has improved, critical components remain severely underreported with significant disparities across disciplines. Our work provides a scalable AI framework to audit and improve RCT reporting, offering actionable insights for journals, researchers, and policymakers to enhance research integrity and clinical translation.