Do large language models meet professional standards in rhinoplasty?? A comparative evaluation with AAO-HNS guidelines
摘要
Large language models are increasingly used by patients to obtain medical information, yet their reliability in addressing specialty-specific surgical questions remains unclear. This study evaluated the performance of publicly available large language models in answering clinical questions about rhinoplasty based on professional guidelines.
MethodsIn this cross-sectional, comparative evaluation, twenty-one questions were selected from the American Academy of Otolaryngology–Head and Neck Surgery Clinical Practice Guidelines for rhinoplasty. Four large language models (ChatGPT version 4.0, ChatGPT version 3.5, Claude version 2.1, and Gemini version 1.5) were evaluated. Four otolaryngologists independently assessed each response for accuracy, extensiveness, misleading information, resource quality, overall reliability, and citation of clinical guidelines. Statistical comparisons were performed using Fisher’s exact test with correction for multiple comparisons. Five lay reviewers independently rated each model’s responses for clarity, usefulness, reassurance, and trustworthiness. Reproducibility was tested across three sessions for a subset of questions.
ResultsAll models achieved high accuracy (≥ 90%) and strong reliability, with excellent inter-rater agreement (Fleiss’ Kappa = 0.93). Significant differences were observed in extensiveness (p = 0.031), misleading content (p = 0.028), resource quality (p < 0.001), and citation frequency (p = 0.004). ChatGPT version 4.0 had the highest citation quality; ChatGPT version 3.5 provided the most extensive responses. Layperson reviewers rated all models highly across all domains.
ConclusionsLarge language models offer accurate, reliable, and patient-friendly responses to rhinoplasty questions, though citation practices and information depth vary. Clinical oversight remains essential to ensure safe use in patient education and decision-making.
Level of evidence: Non ratable.