Utility of Large Language Models for Congenital Microtia Reconstruction Education: Comparison of the Performance of Claude, GPT, and Gemini
摘要
Children with microtia and their parents require comprehensive information to make informed decisions about treatment options.
ObjectiveWe evaluated the effectiveness of various large language models (LLMs) in providing preoperative education for congenital microtia reconstruction (CMR) by analyzing their responses to related inquiries.
MethodsTen plastic surgeons developed 13 CMR-related preoperative education strategies and input 14 text commands into Claude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro during an online session. Five experts evaluated these language model’s responses for correctness, completeness, logic, and potential harm, while five postoperative patients’ parent reviewed the education materials for readability and value. All responses were also analyzed for readability using the context package.
ResultsThe results showed no statistically significant differences among Gemini, Claude, and GPT in the evaluation metrics of accuracy, completeness, and potential risk. In terms of logicality and overall rating, Gemini’s responses were significantly superior to GPT. Preoperative patient education materials generated by GPT received the highest DISCERN scores, significantly outperforming those from Claude and Gemini. From the perspective of patient’s parent’s, there are no statistically significant differences among Gemini, Claude, and GPT. Objective assessments of readability confirmed that Claude’s materials were easier to understand compared to those from the other models.
ConclusionClaude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro effectively addressed patient inquiries and produced clear pre-surgical education materials. However, these LLMs should not be used independently for patient education without expert supervision to ensure accuracy and completeness.
Level of Evidence IVThis journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.