Purpose <p>To evaluate the reliability and accuracy of Large Language Models in answering patient Frequently Asked Questions about adult neck masses.</p> Methods <p>Twenty-four questions from the American Academy of Otolaryngology–Head and Neck Surgery were presented to ChatGPT, Claude, and Gemini. Five independent otolaryngologists evaluated responses using six criteria: accuracy, extensiveness, misleading information, resource quality, guideline citations, and overall reliability. Statistical analysis used Fisher’s exact tests and Fleiss’ Kappa.</p> Results <p>All models showed high reliability (91.7–100%). Paid GPT and Gemini achieved highest accuracy (95.8%). Extensiveness varied significantly (<i>p</i> = 0.012), with Gemini scoring lowest (62.5%). Resource quality ranged from 58.3% (Claude) to 100% (Paid GPT). Guideline citations were highest for GPT models (50%) and lowest for Gemini (16.7%). Misleading information was rare (0-16.7%). Inter-rater reliability was near-perfect across five reviewers (κ = 0.95).</p> Conclusion <p>Large Language Models demonstrate high reliability and accuracy for neck mass patient education, with paid versions showing marginally better performance. While promising as educational tools, variable guideline adherence and occasional misinformation suggest they should complement rather than replace professional medical advice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Are chatbots a reliable source for patient frequently asked questions on neck masses?

  • Sholem Hack,
  • Shibli Alsleibi,
  • Naseem Saleh,
  • Eran E. Alon,
  • Naomi Rabinovics,
  • Eric Remer

摘要

Purpose

To evaluate the reliability and accuracy of Large Language Models in answering patient Frequently Asked Questions about adult neck masses.

Methods

Twenty-four questions from the American Academy of Otolaryngology–Head and Neck Surgery were presented to ChatGPT, Claude, and Gemini. Five independent otolaryngologists evaluated responses using six criteria: accuracy, extensiveness, misleading information, resource quality, guideline citations, and overall reliability. Statistical analysis used Fisher’s exact tests and Fleiss’ Kappa.

Results

All models showed high reliability (91.7–100%). Paid GPT and Gemini achieved highest accuracy (95.8%). Extensiveness varied significantly (p = 0.012), with Gemini scoring lowest (62.5%). Resource quality ranged from 58.3% (Claude) to 100% (Paid GPT). Guideline citations were highest for GPT models (50%) and lowest for Gemini (16.7%). Misleading information was rare (0-16.7%). Inter-rater reliability was near-perfect across five reviewers (κ = 0.95).

Conclusion

Large Language Models demonstrate high reliability and accuracy for neck mass patient education, with paid versions showing marginally better performance. While promising as educational tools, variable guideline adherence and occasional misinformation suggest they should complement rather than replace professional medical advice.