Objective <p>This study assessed the accuracy, comprehensiveness, and consistency of three prominent large language model (LLM)-based chatbots—ChatGPT, Gemini, and Grok—in delivering medical information about Familial Mediterranean Fever (FMF), a rare autoinflammatory disorder.</p> Methods <p>Forty-nine frequently asked, patient-focused questions on FMF were categorized into four domains: basic knowledge, diagnosis, treatment, and recovery/risks/complications/follow-up. Each question was submitted individually to ChatGPT, Gemini, and Grok. Two clinical experts independently assessed the responses using a three-point scale: comprehensive/correct, incomplete/partially correct, and mixed/misleading. Response reproducibility was evaluated by repeating the queries 1&#xa0;week later.</p> Results <p>Gemini achieved the highest accuracy (87.7% comprehensive/correct responses), followed by Grok (83.6%) and ChatGPT (81.6%). ChatGPT demonstrated the greatest consistency (89.7% reproducibility) and provided no misleading content. Grok was the only model to produce misleading responses (4.0%), primarily in diagnosis-related answers. All chatbots performed best in the recovery/follow-up category; diagnostic performance was notably lower.</p> Conclusion <p>LLM-based chatbots demonstrated promising but varied performance in delivering FMF-related medical information. While overall accuracy was high, inconsistencies and occasional misinformation highlight the need for expert supervision, especially for complex clinical topics. These tools may enhance patient education, but should not replace professional guidance. Future research should assess multilingual capabilities and user comprehension to ensure safe and equitable integration into digital health environments.<Table Float="No" ID="Taba"> <tgroup cols="2"> <colspec align="justify" colname="c1" colnum="1" /> <colspec align="justify" colname="c2" colnum="2" /> <tbody> <row> <entry nameend="c2" namest="c1"> <p><b>Key Points</b></p> <p>• <i>This study is the first to systematically compare ChatGPT, Gemini, and Grok regarding their accuracy, consistency, and trustworthiness in providing information on Familial Mediterranean Fever (FMF).</i></p> <p>• <i>Gemini demonstrated the highest accuracy, while ChatGPT showed the best reproducibility and produced no misleading information.</i></p> <p>• <i>The findings highlight the potential and limitations of AI chatbots in patient education for rare autoinflammatory diseases, underscoring the need for expert oversight.</i></p> </entry> </row> </tbody> </tgroup> </Table></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChatGPT, Gemini, and Grok on familial mediterranean fever: are they trustworthy?

  • Selin Cilli Hayıroğlu,
  • Tugce Bozkurt

摘要

Objective

This study assessed the accuracy, comprehensiveness, and consistency of three prominent large language model (LLM)-based chatbots—ChatGPT, Gemini, and Grok—in delivering medical information about Familial Mediterranean Fever (FMF), a rare autoinflammatory disorder.

Methods

Forty-nine frequently asked, patient-focused questions on FMF were categorized into four domains: basic knowledge, diagnosis, treatment, and recovery/risks/complications/follow-up. Each question was submitted individually to ChatGPT, Gemini, and Grok. Two clinical experts independently assessed the responses using a three-point scale: comprehensive/correct, incomplete/partially correct, and mixed/misleading. Response reproducibility was evaluated by repeating the queries 1 week later.

Results

Gemini achieved the highest accuracy (87.7% comprehensive/correct responses), followed by Grok (83.6%) and ChatGPT (81.6%). ChatGPT demonstrated the greatest consistency (89.7% reproducibility) and provided no misleading content. Grok was the only model to produce misleading responses (4.0%), primarily in diagnosis-related answers. All chatbots performed best in the recovery/follow-up category; diagnostic performance was notably lower.

Conclusion

LLM-based chatbots demonstrated promising but varied performance in delivering FMF-related medical information. While overall accuracy was high, inconsistencies and occasional misinformation highlight the need for expert supervision, especially for complex clinical topics. These tools may enhance patient education, but should not replace professional guidance. Future research should assess multilingual capabilities and user comprehension to ensure safe and equitable integration into digital health environments.

Key Points

This study is the first to systematically compare ChatGPT, Gemini, and Grok regarding their accuracy, consistency, and trustworthiness in providing information on Familial Mediterranean Fever (FMF).

Gemini demonstrated the highest accuracy, while ChatGPT showed the best reproducibility and produced no misleading information.

The findings highlight the potential and limitations of AI chatbots in patient education for rare autoinflammatory diseases, underscoring the need for expert oversight.