ChatGPT, Gemini, and Grok on familial mediterranean fever: are they trustworthy?
摘要
This study assessed the accuracy, comprehensiveness, and consistency of three prominent large language model (LLM)-based chatbots—ChatGPT, Gemini, and Grok—in delivering medical information about Familial Mediterranean Fever (FMF), a rare autoinflammatory disorder.
MethodsForty-nine frequently asked, patient-focused questions on FMF were categorized into four domains: basic knowledge, diagnosis, treatment, and recovery/risks/complications/follow-up. Each question was submitted individually to ChatGPT, Gemini, and Grok. Two clinical experts independently assessed the responses using a three-point scale: comprehensive/correct, incomplete/partially correct, and mixed/misleading. Response reproducibility was evaluated by repeating the queries 1 week later.
ResultsGemini achieved the highest accuracy (87.7% comprehensive/correct responses), followed by Grok (83.6%) and ChatGPT (81.6%). ChatGPT demonstrated the greatest consistency (89.7% reproducibility) and provided no misleading content. Grok was the only model to produce misleading responses (4.0%), primarily in diagnosis-related answers. All chatbots performed best in the recovery/follow-up category; diagnostic performance was notably lower.
ConclusionLLM-based chatbots demonstrated promising but varied performance in delivering FMF-related medical information. While overall accuracy was high, inconsistencies and occasional misinformation highlight the need for expert supervision, especially for complex clinical topics. These tools may enhance patient education, but should not replace professional guidance. Future research should assess multilingual capabilities and user comprehension to ensure safe and equitable integration into digital health environments.