Background <p>The present study aimed to evaluate the performance of four large language models (LLMs)—ChatGPT-4o (OpenAI, San Francisco, CA, USA), Gemini 2.5 Flash (Google AI, Mountain View, CA, USA), Claude Sonnet 4 (Anthropic, San Francisco, CA, USA), and DeepSeek V3 (DeepSeek AI, Hangzhou, China)—in answering multiple-choice questions related to maxillofacial prosthetics in both Turkish and English.</p> Methods <p>A total of 45 five-option multiple-choice questions were developed based on <i>Clinical Maxillofacial Prosthetics</i> by Thomas D. Taylor. Each question was submitted to the LLMs in two languages (Turkish and English). Responses were scored by three prosthodontists using a 3-point scale to evaluate both correctness and explanatory quality of the answers. Statistical analyses were performed to evaluate performance differences and cross-lingual consistency. A <i>p</i>-value &lt; 0.05 was considered statistically significant.</p> Results <p>No statistically significant differences were observed among the four LLMs in either language (<i>p</i> = 0.128 for English; <i>p</i> = 0.729 for Turkish). Gemini achieved the highest score in English, with 81.1% accuracy, while Claude and DeepSeek both scored 78.9% in Turkish. Strong and statistically significant positive correlations were observed between English and Turkish scores across all LLMs, indicating consistent relative performance regardless of language.</p> Conclusions <p>When answering maxillofacial prosthetics questions, LLMs demonstrated comparable performance in both Turkish and English. Their consistent ranking suggests potential reliability in cross-lingual knowledge generation and highlights their value as effective tools in multilingual dental education and clinical decision support.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-lingual performance of large language models in maxillofacial prosthodontics: a comparative evaluation

  • Irem Sozen Yanik,
  • Dilara Sahin Hazir,
  • Damla Bilgin Avsar

摘要

Background

The present study aimed to evaluate the performance of four large language models (LLMs)—ChatGPT-4o (OpenAI, San Francisco, CA, USA), Gemini 2.5 Flash (Google AI, Mountain View, CA, USA), Claude Sonnet 4 (Anthropic, San Francisco, CA, USA), and DeepSeek V3 (DeepSeek AI, Hangzhou, China)—in answering multiple-choice questions related to maxillofacial prosthetics in both Turkish and English.

Methods

A total of 45 five-option multiple-choice questions were developed based on Clinical Maxillofacial Prosthetics by Thomas D. Taylor. Each question was submitted to the LLMs in two languages (Turkish and English). Responses were scored by three prosthodontists using a 3-point scale to evaluate both correctness and explanatory quality of the answers. Statistical analyses were performed to evaluate performance differences and cross-lingual consistency. A p-value < 0.05 was considered statistically significant.

Results

No statistically significant differences were observed among the four LLMs in either language (p = 0.128 for English; p = 0.729 for Turkish). Gemini achieved the highest score in English, with 81.1% accuracy, while Claude and DeepSeek both scored 78.9% in Turkish. Strong and statistically significant positive correlations were observed between English and Turkish scores across all LLMs, indicating consistent relative performance regardless of language.

Conclusions

When answering maxillofacial prosthetics questions, LLMs demonstrated comparable performance in both Turkish and English. Their consistent ranking suggests potential reliability in cross-lingual knowledge generation and highlights their value as effective tools in multilingual dental education and clinical decision support.