Evaluation of artificial ıntelligence use in ankylosing spondylitis with ChatGPT-4: patient and physician perspectives
摘要
This study aims to evaluate the accuracy and comprehensiveness of the information provided by ChatGPT-4, an artificial intelligence-based system, regarding ankylosing spondylitis (AS) from the perspectives of patients and physicians.
MethodIn this cross-sectional study, 75 questions were asked of ChatGPT-4. These were the most frequently asked questions about AS on Google Trends (group 1), and questions derived from ASAS/EULAR recommendations (group 2 and group 3). Group 2 consisted of open-ended questions, and group 3 consisted of case questions. Two expert rheumatologists scored the responses for accuracy and comprehensiveness. A six-point Likert scale was used to assess accuracy, and a three-point scale for completeness.
ResultsThe accuracy and completeness scores analyzed in this study were found to be 5.32 ± 1.4 and 2.76 ± 0.5 for group 1, 5.36 ± 1.1 and 2.72 ± 0.45 for group 2, and 4.24 ± 1.96 and 2.36 ± 0.63 for group 3, respectively. There was a significant difference in accuracy and completeness scores between the groups (p = 0.044 and p = 0.019). Cohen’s kappa coefficient showed excellent agreement with values of 0.88 for accuracy and 0.90 for completeness.
ConclusionWhile the responses to questions in groups 1 and 2 were satisfactory in terms of accuracy and comprehensiveness, the responses to complex case questions in group 3 were not sufficient. ChatGPT-4 appears to be a useful resource for patient education, but response mechanisms for complex clinical scenarios need to be improved. Furthermore, the potential for generating false or fabricated data should be considered; therefore, physicians should evaluate and verify responses.