An automated approach to improve clinical trial registration and to identify outcome changes on ClinicalTrials.gov
摘要
It is difficult to ensure clinical trial outcomes are defined completely in registrations, and to check outcome changes between registration and posting results. We developed a large language model (LLM)-based approach to evaluate outcome definitions and changes on ClinicalTrials.gov. Our LLM-based approach accurately identified incomplete outcomes in prospective trial registrations (sensitivity, 0.91 [95% confidence interval {CI}, 0.87–0.94]; positive predictive value [PPV], 0.98 [95% CI, 0.97–1.00]). Comparing prospective registrations with posted results, it correctly identified 96.1% of missing and 97.9% of added outcomes. It identified all outcomes with a priority change (e.g., from primary to secondary) and 99.2% without one. Sensitivity and PPV for identifying any outcome change were 0.97 (95% CI, 0.95–0.99) and 0.95 (95% CI, 0.89–0.98), respectively. Estimated costs per trial were $0.13 for o3-mini, $0.30 for GPT-4o, and $1.80 for o1. The accuracy of our approach suggests it could be used to improve registration quality and to detect outcome changes at large scale and at low cost.