Comparative evaluation of AI platforms “Google Gemini 2.5 Flash, Google Gemini 2.0 Flash, DeepSeek V3 and ChatGPT 4o” in solving multiple-choice questions from different subtopics of anatomy
摘要
The rise of artificial intelligence (AI) based large language models (LLMs) had a profound impact on medical education. Given the widespread use of multiple-choice questions (MCQs) in anatomy education, it is likely that such queries are commonly directed to AI tools. The current study compared the accuracy level of different AI platforms for solving MCQs from various subtopics in Anatomy.
MethodsA total of 55 MCQs from different subtopics of Anatomy were enquired using Google Gemini 2.0 Flash, Google Gemini Flash 2.5 DeepSeek V3 and ChatGPT 4o AI platforms. The accuracy rate was calculated. Chi-square test and Fisher’s exact test was employed for statistical analysis.
ResultsOverall accuracy of 95.9% was observed across all platforms. When ranked by performance, Google Gemini 2.5 Flash performed the best, followed by Google Gemini 2.0, with Chat GPT 4o and DeepSeek V3. No statistically significant performance difference among different AI models was observed. General Anatomy was identified as the most challenging area across all the models.
ConclusionsAll models performed exceptionally well on anatomy-focused MCQs. Google Gemini models show superior overall performance. However, certain errors persist that cannot be overlooked, highlighting the continued need for human oversight and expert validation. As per best available literature, this is the first study to include MCQS from different anatomy subtopics and to compare the performance of DeepSeek and Google Gemini Flash 2.5 for the task.