Evaluation of the effectiveness of ChatGPT in supporting the management of autoimmune hepatitis effectiveness of ChatGPT in autoimmune hepatitis
摘要
This study aims to explore ChatGPT’s ability to provide comprehensive information on AIH, its potential role in supporting the diagnostic and therapeutic processes for the disease and its broader implications for the use of artificial intelligence in healthcare.
Materials and methodsA total of 45 questions were designed and organized into 3 groups, each consisting of 15 questions (Group 1, Group 2, Group 3). The questions were categorized based on diffuculty, with Group 1 being the easiest and Group 3 the most challenging. Additionally, 5 case-based questions were formulated and ChatGPT was asked to provide diagnoses for these cases. All questions were re-asked after 14 days to evaluate the short-term stability and consistency of the artificial intelligence responses.
ResultsThere was no statistically significant difference in the responses provided by ChatGPT initially (p = 0.328). However, when the questions were re-asked after 14 days, a significant difference was observed between the groups in terms of accuracy and completeness (p = 0.045 and p = 0.015, respectively). This difference was primarily due to the lower accuracy level in the responses for group 3 questions. Furthermore, when comparing ChatGPT’s responses at 2-week intervals, a statistically significant improvement was noted, except for the completeness score of group 3 questions and the case-based questions.
ConclusionOur findings suggest that ChatGPT may play a greater role in patient education and in supporting physicians in disease management. Future studies that test AI-generated responses from multiple perspectives and involve more centers could provide clearer guidance on this topic.