<p>Manual referral triage in tertiary hospitals is resource-intensive and prone to inconsistency. This study evaluates the feasibility of an on-premises large language model (LLM), using a roster-embedded prompting strategy, to assist in subspecialty triage of referral letters. Utilizing a dataset of 6624 electronic referral letters from Samsung Medical Center, we deployed an open-source LLM (Qwen-2.5-32B) within a secure infrastructure, incorporating real-time clinician availability. In a hold-out test set (<i>n</i> = 680), the LLM achieved a baseline accuracy of 75.4% (95% CI, 72.2–78.7%) compared to human coordinators, which improved to 84.7% (95% CI, 81.9–87.4%) after expert adjudication of discordant cases. Performance of frequently referred subspecialties was accuracy of 86.1%, and performance of infrequent ones was 74.4%, which were adjudicated. Error analysis revealed that misclassifications primarily occurred between clinically adjacent departments rather than at random, and a small subset of referrals (5.9%) lacked sufficient information for unique assignment. These findings demonstrate that our roster-embedded LLM can provide highly valid subspecialty assignments and serve a complementary role in referral workflows. By integrating such models into human-in-the-loop systems, tertiary hospitals may significantly enhance operational efficiency and reduce administrative burdens.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large language model-assisted referral triage automation in a tertiary hospital

  • Byeongjae Kang,
  • Minjun Son,
  • Wonoh Jeong,
  • Seo Yeong Jeong,
  • Young Joo Kim,
  • Hye-Eun Park,
  • Hyun Kyung Lee,
  • Myung Jin Chung,
  • Taeyoung Kim,
  • Kwangmo Yang

摘要

Manual referral triage in tertiary hospitals is resource-intensive and prone to inconsistency. This study evaluates the feasibility of an on-premises large language model (LLM), using a roster-embedded prompting strategy, to assist in subspecialty triage of referral letters. Utilizing a dataset of 6624 electronic referral letters from Samsung Medical Center, we deployed an open-source LLM (Qwen-2.5-32B) within a secure infrastructure, incorporating real-time clinician availability. In a hold-out test set (n = 680), the LLM achieved a baseline accuracy of 75.4% (95% CI, 72.2–78.7%) compared to human coordinators, which improved to 84.7% (95% CI, 81.9–87.4%) after expert adjudication of discordant cases. Performance of frequently referred subspecialties was accuracy of 86.1%, and performance of infrequent ones was 74.4%, which were adjudicated. Error analysis revealed that misclassifications primarily occurred between clinically adjacent departments rather than at random, and a small subset of referrals (5.9%) lacked sufficient information for unique assignment. These findings demonstrate that our roster-embedded LLM can provide highly valid subspecialty assignments and serve a complementary role in referral workflows. By integrating such models into human-in-the-loop systems, tertiary hospitals may significantly enhance operational efficiency and reduce administrative burdens.