Performance of artificial intelligence chatbots in providing feeding management and oral health guidance for children with cleft lip and palate: ChatGPT 5.2 vs Gemini 3 Pro
摘要
This study aims to comparatively evaluate the quality, reliability, and readability of responses generated by ChatGPT 5.2 and Gemini 3 Pro regarding feeding management and oral health guidance for children with cleft lip and palate (CLP). A cross-sectional design was used. Both models were asked 20 questions on feeding and oral-dental care in infants and children with CLP. Response quality was assessed using the Global Quality Score (GQS) and the CLEAR tool, reliability with the modified-DISCERN (m-DISCERN), and readability with the Flesch Reading Ease (FRES) and Flesch–Kincaid Grade Level (FKGL). The mean GQS was significantly higher for Gemini 3 Pro than for ChatGPT 5.2 (4.51 ± 0.29 vs. 3.76 ± 0.33; p < 0.001). The CLEAR tool score was likewise significantly higher in Gemini 3 Pro compared with ChatGPT 5.2 (22.95 ± 0.94 vs. 19.40 ± 1.47; p < 0.001), and the m-DISCERN scores were also significantly higher for Gemini 3 Pro (4.0 [4.0–4.0]) than for ChatGPT 5.2 (3.0 [2.0–3.0]; p < 0.001). The FRES values were significantly higher for Gemini 3 Pro compared with ChatGPT 5.2 (49.50 ± 10.33 vs. 23.60 ± 14.48; p < 0.001), whereas the FKGL values were significantly lower (9.77 ± 1.54 vs. 13.44 ± 2.81; p < 0.001). Correlation analysis revealed strong positive correlations between GQS and CLEAR (r = 0.753) and between GQS and m-DISCERN (r = 0.740), while FKGL showed significant negative correlations with all quality and reliability measures (−0.566≤ r ≤ −0.484; all p < 0.001).
Conclusion: Gemini 3 Pro outperformed ChatGPT 5.2 across content quality, reliability, and readability. Although both models can support guidance in CLP care, AI chatbot outputs should be used as complementary tools alongside professional clinical guidance.