“Comparative analysis of large language models against the NHS 111 online triaging for emergency ophthalmology”
摘要
This study presents a comprehensive evaluation of the performance of various large language models in generating responses for ophthalmology emergencies and compares their accuracy with the established United Kingdom’s National Health Service 111 online system.
MethodsWe included 21 ophthalmology-related emergency scenario questions from the NHS 111 triaging algorithm. These questions were based on four different ophthalmology emergency themes as laid out in the NHS 111 algorithm. Responses generated from NHS 111 online, were compared to different LLM-chatbots responses to determine the accuracy of LLM responses. We included a range of models including ChatGPT-3.5, Google Bard, Bing Chat, and ChatGPT-4.0. The accuracy of each LLM-chatbot response was compared against the NHS 111 Triage using a two-prompt strategy. Answers were graded as following: −2 graded as “Very poor”, −1 as “Poor”, O as “No response”, 1 as “Good”, 2 as “Very good” and 3 graded as “Excellent”.
ResultsOverall LLMs’ attained a good accuracy in this study compared against the NHS 111 responses. The score of ≥1 graded as “Good” was achieved by 93% responses of all LLMs. This refers to at least part of this answer having correct information as well as absence of any wrong information. There was no marked difference and very similar results seen overall on both prompts.
ConclusionsThe high accuracy and safety observed in LLM responses support their potential as effective tools for providing timely information and guidance to patients. LLMs hold promise in enhancing patient care and healthcare accessibility in digital age.