Evaluation of large language models in patient education for hyperthyroidism: A comparative study of chatgpt, gemini, and deepseek
摘要
Hyperthyroidism is a common endocrine disorder that requires long-term management, heavily relying on effective patient education. This study aims to evaluate the application value of three mainstream large language models (LLMs) ChatGPT, Gemini, and DeepSeek in the education of hyperthyroidism patients.
MethodsWe developed a standardized question bank containing 20 issues related to hyperthyroidism. The three LLMs were prompted to generate responses for each question. Five endocrinology experts performed a double-blind evaluation using a Likert 5-point scale across five dimensions: relevance, accuracy, comprehensibility, comprehensiveness, and humanistic care. Additionally, the Flesch-Kincaid readability formula was used to analyze text complexity.
ResultsThe scores of the three large language models demonstrated statistically significant differences (p < 0.05). DeepSeek achieved the highest scores in relevance (4.80 ± 0.45), accuracy (4.59 ± 0.57), and comprehensiveness (4.70 ± 0.46). Gemini excelled in comprehensibility (4.49 ± 0.58) and humanistic care (4.27 ± 0.33). Notably, The Flesch Reading Ease Index classified the text generated by all three LLMs as ‘Difficult’ to read (DeepSeek: 38, Gemini: 38, ChatGPT: 37).
ConclusionsLLMs show potential in hyperthyroidism patient education, but there is still room for improvement across various dimensions. Patients and healthcare professionals should consider these models as supplementary tools rather than replacements for professional medical personnel. Further research is needed to explore the clinical application of artificial intelligence in healthcare.