Incomplete reporting of a study’s methods and results hinders efforts to evaluate and reproduce research findings in randomized controlled trials (RCTs), leading to potential harm. CONSORT, a set of widely endorsed RCT reporting guidelines, was designed to mitigate such harm by ensuring transparency, reproducibility and safety in RCTs. Evaluating adherence to CONSORT requires a level of language understanding that has previously precluded broad systematic assessment. However, recent advances in large language models (LLMs) now make it possible to evaluate the quality of RCTs at scale. We demonstrate GPT-4o-mini, used out of the box to achieve state-of-the-art performance in evaluating RCT quality (F1 score: 0.85; precision: 0.96), with results validated by expert human annotators showing 92.24% agreement across 50 papers. Applying this tool to 21,041 open-access RCTs (1966–2024), we reveal temporal and domain trends: while overall adherence to CONSORT has improved, critical components remain severely underreported with significant disparities across disciplines. Our work provides a scalable AI framework to audit and improve RCT reporting, offering actionable insights for journals, researchers, and policymakers to enhance research integrity and clinical translation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Reporting Quality of 21,041 Randomized Controlled Trial Articles with Large Language Models: A Large-Scale Transparency Analysis

  • Apoorva Srinivasan,
  • Sophia Kivelson,
  • Nadine A. Friedrich,
  • Jacob Berkowitz,
  • Nicholas Tatonetti

摘要

Incomplete reporting of a study’s methods and results hinders efforts to evaluate and reproduce research findings in randomized controlled trials (RCTs), leading to potential harm. CONSORT, a set of widely endorsed RCT reporting guidelines, was designed to mitigate such harm by ensuring transparency, reproducibility and safety in RCTs. Evaluating adherence to CONSORT requires a level of language understanding that has previously precluded broad systematic assessment. However, recent advances in large language models (LLMs) now make it possible to evaluate the quality of RCTs at scale. We demonstrate GPT-4o-mini, used out of the box to achieve state-of-the-art performance in evaluating RCT quality (F1 score: 0.85; precision: 0.96), with results validated by expert human annotators showing 92.24% agreement across 50 papers. Applying this tool to 21,041 open-access RCTs (1966–2024), we reveal temporal and domain trends: while overall adherence to CONSORT has improved, critical components remain severely underreported with significant disparities across disciplines. Our work provides a scalable AI framework to audit and improve RCT reporting, offering actionable insights for journals, researchers, and policymakers to enhance research integrity and clinical translation.