Background <p>Large language models (LLMs) demonstrate increasing potential in healthcare applications, yet their clinical utility in specialized pediatric medicine remains inadequately characterized. This study evaluated LLM performance in pediatric urology to establish evidence-based implementation frameworks.</p> Methods <p>We conducted a two-phase evaluation of thirteen current-generation LLMs between January and April 2025. Phase 1 assessed clinical decision support using 180 standardized cases stratified by complexity (126 routine, 54 complex) based on European Association of Urology pediatric guidelines. Phase 2 evaluated patient education effectiveness across four common pediatric urological conditions. Performance was measured using mathematical similarity metrics and standardized expert evaluation by seven board-certified pediatric urologists employing a validated 100-point scoring system.</p> Results <p>LLMs demonstrated superior performance in routine clinical scenarios compared to complex cases (76% vs. 41% accuracy, respectively). OpenAI’s GPT-4o3 achieved the highest performance in complex cases (61.4 ± 3.1 vs. 51.8 ± 3.7 for other models, <i>p</i> &lt; 0.001), with significantly lower hallucination rates (0.05 ± 0.02 vs. 0.18 ± 0.04, <i>p</i> &lt; 0.001). Patient education materials showed strong content alignment (78% similarity with reference standards) and appropriate readability levels. Diagnostic confidence demonstrated strong correlation with actual performance (<i>r</i> = 0.82, <i>p</i> &lt; 0.001).</p> Conclusions <p>LLMs show promise as supportive tools in pediatric urology, particularly for patient education and routine clinical scenarios. However, significant limitations in complex case management necessitate careful implementation with mandatory physician oversight. A tiered approach prioritizing patient education while restricting complex clinical decision-making represents the most appropriate implementation strategy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From digital assistants to clinical partners: revolutionizing pediatric urology through large language model-powered decision support and patient education

  • Mert Başaranoğlu,
  • Erdem Akbay,
  • Erim Erdem

摘要

Background

Large language models (LLMs) demonstrate increasing potential in healthcare applications, yet their clinical utility in specialized pediatric medicine remains inadequately characterized. This study evaluated LLM performance in pediatric urology to establish evidence-based implementation frameworks.

Methods

We conducted a two-phase evaluation of thirteen current-generation LLMs between January and April 2025. Phase 1 assessed clinical decision support using 180 standardized cases stratified by complexity (126 routine, 54 complex) based on European Association of Urology pediatric guidelines. Phase 2 evaluated patient education effectiveness across four common pediatric urological conditions. Performance was measured using mathematical similarity metrics and standardized expert evaluation by seven board-certified pediatric urologists employing a validated 100-point scoring system.

Results

LLMs demonstrated superior performance in routine clinical scenarios compared to complex cases (76% vs. 41% accuracy, respectively). OpenAI’s GPT-4o3 achieved the highest performance in complex cases (61.4 ± 3.1 vs. 51.8 ± 3.7 for other models, p < 0.001), with significantly lower hallucination rates (0.05 ± 0.02 vs. 0.18 ± 0.04, p < 0.001). Patient education materials showed strong content alignment (78% similarity with reference standards) and appropriate readability levels. Diagnostic confidence demonstrated strong correlation with actual performance (r = 0.82, p < 0.001).

Conclusions

LLMs show promise as supportive tools in pediatric urology, particularly for patient education and routine clinical scenarios. However, significant limitations in complex case management necessitate careful implementation with mandatory physician oversight. A tiered approach prioritizing patient education while restricting complex clinical decision-making represents the most appropriate implementation strategy.