Comparison of ChatGPT, Gemini, and DeepSeek responses to questions on medication-related osteonecrosis of the jaw (MRONJ): a comparative analysis of expert-rated clinical quality, readability, response time, and clinical safety risk
摘要
Medication-related osteonecrosis of the jaw (MRONJ) is a clinically significant complication associated with antiresorptive and antiangiogenic therapies, and its prevention and management depend on reliable, guideline-concordant information. Large language models (LLMs) are increasingly used to generate medical information, but their performance in MRONJ remains insufficiently investigated. This study compared the responses generated by ChatGPT- 5.4 Thinking, Gemini 3 Deep Think, and DeepSeek-V3.2 in thinking mode to MRONJ-related open-ended questions in terms of clinical quality, readability, and response time.
MethodsIn this cross-sectional comparative study, a standardized 30-item question set was developed from the MASCC/ISOO/ASCO Clinical Practice Guideline and the Italian Consensus Update. Each question was entered separately into each model under standardized conditions, yielding 90 responses. Outputs were anonymized and independently evaluated by three blinded oral and maxillofacial surgeons using a modified Global Quality Scale on a five-point Likert scale. A post hoc Clinical Risk Classification was additionally performed to distinguish minor informational deficiencies from omissions or inaccuracies with greater potential clinical consequences. Readability was assessed using five established English-language indices, and response time was measured with an online digital stopwatch. Appropriate statistical tests were applied, and inter-evaluator reliability was assessed using the intraclass correlation coefficient.
ResultsInter-evaluator agreement was high across all models (all p < 0.001). Mean expert quality scores differed significantly among models (p < 0.001). Post hoc comparisons showed that ChatGPT-5.4 Thinking received significantly lower expert ratings than both DeepSeek-V3.2 and Gemini 3 Deep Think, whereas no statistically significant difference was detected between DeepSeek-V3.2 and Gemini 3 Deep Think. Response time also differed significantly (p < 0.001): DeepSeek-V3.2 produced the fastest responses, Gemini 3 Deep Think showed intermediate latency, and ChatGPT-5.4 Thinking had the longest response times. Readability analyses revealed significant inter-model differences across all indices, with Gemini 3 Deep Think generating the most complex outputs and DeepSeek-V3.2 producing comparatively more readable responses. Clinical Risk Classification showed that most responses were categorized as having no clinically relevant risk or low clinical risk, while moderate-risk classifications were uncommon and no high-risk responses were identified.
ConclusionsUnder single-prompt, single-session benchmark conditions, potentially clinically relevant differences were observed among the models in MRONJ-related response quality, readability, response time, and clinical risk classification. DeepSeek-V3.2 and Gemini 3 Deep Think achieved higher expert-rated quality scores than ChatGPT-5.4 Thinking, while DeepSeek-V3.2 also showed advantages in speed and readability. These findings do not indicate general model superiority and support using LLMs only as adjunctive informational tools, not substitutes for specialist and guideline-based clinical decision-making.