COGnitive Network Evaluation Toolkit for Medical Domains: Evaluating Large Language Model Performance in Medical Diagnostics—An Assessment Framework and Dataset for Healthcare AI
摘要
We introduce Cognitive Network Evaluation Toolkit for Medical Domains (COGNET-MD), which constitutes a novel benchmark for evaluation of the correctness of Large Language Model responses in the medical domain. Specifically, COGNET-MD consists of a dataset of Multiple Choice Quizzes (MCQs) along with a proposed framework for scoring the LLM responses. The MCQs have been curated by medical experts and are characterized by a varying degree of difficulty and a number of correct choices which may range from only one to several per question. Because of the latter, the proposed evaluation framework may award partial or full credit depending on the number of correct choices returned by the LLM, while it deducts half a point for each erroneous choice. The current first version (1.0) of the MCQ dataset includes questions from the medical domains of Psychiatry, Dentistry, Pulmonology, Dermatology and Endocrinology. We employed the COGNET-MD benchmark using GPT-4 for inference from this dataset. We found that GPT-4 performed well regarding questions with only one correct answer, as it achieved full credit for 64/75 questions in Psychiatry, 26/36 in Pulmonology and 53/76 in Dentistry. However, when penalized for erroneous choices, GPT-4 scored averagely, achieving only a 55% accuracy.