<p>Early identification of postoperative delirium (POD) remains challenging. This retrospective observational study compared the performance of large language models (LLMs), Llama-3-70B and GPT-4o, and physicians in predicting clinically significant POD, defined as either requiring antipsychotics or diagnosis of delirium by neurologists following consultation for delirium-related symptoms. The c-statistics of Llama-3-70B and GPT-4o were 0.74 and 0.76, respectively. LLMs showed higher sensitivity (Llama-3-70B, 0.900; GPT-4o, 0.868; physicians, 0.723) and lower specificity (0.463, 0.547, and 0.814, respectively) than physicians. Inter-rater agreement was almost perfect for both Llama-3-70B and GPT-4o (Fleiss’ kappa = 0.852 and 0.854, respectively) but fair for physicians (0.219). Both LLMs detected clinically significant POD approximately one day earlier than physicians (Kaplan-Meier analysis, median time to diagnosis: Llama-3-70B, 34.5 h; GPT-4o, 37.5 h; physicians, 62.9 h; log-rank <i>P</i> &lt; 0.001). The integration of LLMs as a complementary screening tool under physician supervision may improve the early, reproducible diagnosis of clinically significant POD.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficacy of large language models in detecting postoperative delirium from unstructured clinical notes: A retrospective cohort study

  • Da-In Eun,
  • Hyung-Chul Lee,
  • Gang Heo,
  • Jae Hoon Jeong,
  • Jipyeong Lee,
  • Hyeonhoon Lee,
  • Sang-Youn Park,
  • Karam Nam,
  • Ho-Jin Lee,
  • Jae-Woo Ju

摘要

Early identification of postoperative delirium (POD) remains challenging. This retrospective observational study compared the performance of large language models (LLMs), Llama-3-70B and GPT-4o, and physicians in predicting clinically significant POD, defined as either requiring antipsychotics or diagnosis of delirium by neurologists following consultation for delirium-related symptoms. The c-statistics of Llama-3-70B and GPT-4o were 0.74 and 0.76, respectively. LLMs showed higher sensitivity (Llama-3-70B, 0.900; GPT-4o, 0.868; physicians, 0.723) and lower specificity (0.463, 0.547, and 0.814, respectively) than physicians. Inter-rater agreement was almost perfect for both Llama-3-70B and GPT-4o (Fleiss’ kappa = 0.852 and 0.854, respectively) but fair for physicians (0.219). Both LLMs detected clinically significant POD approximately one day earlier than physicians (Kaplan-Meier analysis, median time to diagnosis: Llama-3-70B, 34.5 h; GPT-4o, 37.5 h; physicians, 62.9 h; log-rank P < 0.001). The integration of LLMs as a complementary screening tool under physician supervision may improve the early, reproducible diagnosis of clinically significant POD.