<p>Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce <Emphasis FontCategory="NonProportional">TrialPanorama</Emphasis>, a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on <Emphasis FontCategory="NonProportional">TrialPanorama</Emphasis> using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing <Emphasis FontCategory="NonProportional">TrialPanorama</Emphasis> as a foundation for scaling AI in clinical research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking and developing large language models using one million clinical trials

  • Zifeng Wang,
  • Jiacheng Lin,
  • Qiao Jin,
  • Junyi Gao,
  • Jathurshan Pradeepkumar,
  • Pengcheng Jiang,
  • Zhiyong Lu,
  • Jimeng Sun

摘要

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce TrialPanorama, a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on TrialPanorama using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing TrialPanorama as a foundation for scaling AI in clinical research.