Purpose <p>To evaluate and compare the clinical accuracy, reliability, comprehensiveness and readability of three prominent Large Language Models (LLMs) (ChatGPT, Gemini, and Copilot) in the management of pediatric ureteropelvic junction obstruction.</p> Methods <p>A total of 125 unique clinical scenarios representing various stages and complexities of pediatric ureteropelvic junction obstruction(UPJO) were developed. Responses from ChatGPT (GPT-5.3 version, OpenAI), Gemini(Google), and Copilot GPT 5 (Microsoft) were independently evaluated by two expert pediatric urologists across four domains: Clinical Accuracy, Completeness, Reliability, and Readability (scored 1–5). Initial inter-rater discrepancies were resolved through a consensus-building process. Statistical analysis included Kruskal-Wallis tests for performance comparison and weighted Cohen’s kappa for inter-rater agreement.</p> Results <p>ChatGPT demonstrated significantly higher scores in clinical accuracy (4.13 +/- 1.31) and reliability (4.08 +/- 1.30) compared to its counterparts, showing the closest alignment with international guidelines. Conversely, Copilot achieved the highest readability score (4.13 +/- 0.65) but exhibited a “readability-accuracy paradox,” where professional formatting masked frequent clinical inaccuracies (3.49 +/- 1.67). Gemini provided comprehensive content but was hindered by structural deficits and the lowest readability score (2.88 ± 0.83). The expert consensus process successfully improved inter-rater agreement (κ = 0.532) to (κ = 0.586).</p> Conclusion <p>While LLMs show promise as decision-support tools in pediatric surgery, their performance is inconsistent. ChatGPT is currently the most robust model for guideline-based management of UPJO. However, the “deceptive confidence” of models like Copilot poses a risk of misinformation. Future integration should explore multimodal capabilities, including the analysis of imaging and ongoing validation against standardized reporting frameworks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios

  • İbrahim Halil Baloğlu,
  • Gökçe Karlı,
  • Ali Emre Çekmece,
  • Mehmet Özay Özgür,
  • Ahmet Tevfik Albayrak,
  • Kadir Cem Günay,
  • Mustafa Anıl Kılıç,
  • Pınar Zeytun Baloğlu,
  • Kaya Horasanlı

摘要

Purpose

To evaluate and compare the clinical accuracy, reliability, comprehensiveness and readability of three prominent Large Language Models (LLMs) (ChatGPT, Gemini, and Copilot) in the management of pediatric ureteropelvic junction obstruction.

Methods

A total of 125 unique clinical scenarios representing various stages and complexities of pediatric ureteropelvic junction obstruction(UPJO) were developed. Responses from ChatGPT (GPT-5.3 version, OpenAI), Gemini(Google), and Copilot GPT 5 (Microsoft) were independently evaluated by two expert pediatric urologists across four domains: Clinical Accuracy, Completeness, Reliability, and Readability (scored 1–5). Initial inter-rater discrepancies were resolved through a consensus-building process. Statistical analysis included Kruskal-Wallis tests for performance comparison and weighted Cohen’s kappa for inter-rater agreement.

Results

ChatGPT demonstrated significantly higher scores in clinical accuracy (4.13 +/- 1.31) and reliability (4.08 +/- 1.30) compared to its counterparts, showing the closest alignment with international guidelines. Conversely, Copilot achieved the highest readability score (4.13 +/- 0.65) but exhibited a “readability-accuracy paradox,” where professional formatting masked frequent clinical inaccuracies (3.49 +/- 1.67). Gemini provided comprehensive content but was hindered by structural deficits and the lowest readability score (2.88 ± 0.83). The expert consensus process successfully improved inter-rater agreement (κ = 0.532) to (κ = 0.586).

Conclusion

While LLMs show promise as decision-support tools in pediatric surgery, their performance is inconsistent. ChatGPT is currently the most robust model for guideline-based management of UPJO. However, the “deceptive confidence” of models like Copilot poses a risk of misinformation. Future integration should explore multimodal capabilities, including the analysis of imaging and ongoing validation against standardized reporting frameworks.