Cross-lingual performance of large language models in maxillofacial prosthodontics: a comparative evaluation
摘要
The present study aimed to evaluate the performance of four large language models (LLMs)—ChatGPT-4o (OpenAI, San Francisco, CA, USA), Gemini 2.5 Flash (Google AI, Mountain View, CA, USA), Claude Sonnet 4 (Anthropic, San Francisco, CA, USA), and DeepSeek V3 (DeepSeek AI, Hangzhou, China)—in answering multiple-choice questions related to maxillofacial prosthetics in both Turkish and English.
MethodsA total of 45 five-option multiple-choice questions were developed based on Clinical Maxillofacial Prosthetics by Thomas D. Taylor. Each question was submitted to the LLMs in two languages (Turkish and English). Responses were scored by three prosthodontists using a 3-point scale to evaluate both correctness and explanatory quality of the answers. Statistical analyses were performed to evaluate performance differences and cross-lingual consistency. A p-value < 0.05 was considered statistically significant.
ResultsNo statistically significant differences were observed among the four LLMs in either language (p = 0.128 for English; p = 0.729 for Turkish). Gemini achieved the highest score in English, with 81.1% accuracy, while Claude and DeepSeek both scored 78.9% in Turkish. Strong and statistically significant positive correlations were observed between English and Turkish scores across all LLMs, indicating consistent relative performance regardless of language.
ConclusionsWhen answering maxillofacial prosthetics questions, LLMs demonstrated comparable performance in both Turkish and English. Their consistent ranking suggests potential reliability in cross-lingual knowledge generation and highlights their value as effective tools in multilingual dental education and clinical decision support.