Does ChatGPT update itself? Accuracy of ChatGPT in tympanostomy tube guidance: A comparative analysis with current literature
摘要
This study aims to evaluate the accuracy of ChatGPT-4.0 in providing information on tympanostomy tube indications in children, comparing its responses with established clinical guidelines and examining its ability to update itself over time.
MethodsSixteen clinical scenarios from the American Academy of Otolaryngology–Head and Neck Surgery Foundation (AAO-HNSF) guidelines were assessed using 18 specific questions. Responses were evaluated by two otolaryngologists and ChatGPT itself. The final validation was conducted by a senior otolaryngologist. Cohen’s Kappa analysis was performed to assess inter-rater reliability.
ResultsChatGPT-4.0 correctly answered 15.5 out of 16 scenarios (96.8%). The second-stage question of scenario 7 was evaluated as incorrect. When current literature was referenced, all responses reached 100% accuracy. Among the correct answers, 4 scenarios were not fully aligned with the guidelines. However, when responses were based on current literature, all of these answers were found to be fully compliant. The agreement among the three evaluators was perfect, as confirmed by Cohen’s Kappa analysis.
Despite using an updated version (ChatGPT-4.0) and over a year having passed, it was observed that ChatGPT-3.5 answered a previously incorrect scenario in the same incorrect manner. This suggests that the model may have limited capacity for self-updating over time. These findings are consistent with previous research, indicating that ChatGPT provides highly accurate responses regarding tympanostomy tube placement and largely aligns with existing guidelines.
ConclusionChatGPT-4.0 demonstrates high accuracy in providing guideline-based medical information, but its ability to update itself over time appears to be limited. However, when prompted to reference current literature, its accuracy improves significantly. These findings highlight the importance of structured prompting and critical evaluation of AI-generated medical guidance.