Background <p>Although AI (artificial intelligence) chatbot systems are increasingly used to obtain dental information, limited evidence exists regarding the reliability and readability of their responses in laminate veneer–related patient inquiries. The aim of the study is to assess the reliability, quality, and readability of the responses given by 3 different artificial intelligence chatbot systems to frequently asked questions about laminate veneer restorations by patients.</p> Methods <p>Twenty-five frequently asked questions about laminate veneer restorations submitted to ChatGPT 5.2, Gemini 3, and DeepSeek-V3.2. AI chatbots to generate answers. The quality of these answers was assessed by two different experts using GQS (Global Quality Scale), reliability using modified DISCERN, and readability using FRES (Flesch Reading Ease Score) and FKGL (Flesch–Kincaid Grade Level). Normality was assessed with the Shapiro–Wilk test. Depending on data distribution, repeated-measures ANOVA or the Friedman test was used for comparisons among chatbots. Dunn’s test was applied for post hoc analyses. A <i>p</i>-value &lt; 0.05 was considered statistically significant.</p> Results <p>No statistically significant differences were found among chatbot systems in terms of DISCERN, FRES, or FKGL scores (<i>P</i> &gt; 0.05). Interobserver agreement for DISCERN was generally poor, with only Gemini demonstrating statistically significant agreement (ICC = 0.393; <i>P</i> = 0.023). Although a significant overall difference was observed in GQS scores (<i>P</i> = 0.015), no significant pairwise differences were detected. Readability analysis revealed that all chatbot responses exceeded the recommended 6th–8th grade reading level for patient education materials, with mean FKGL values corresponding to approximately 1–2 grade levels above the recommended upper limit. In addition, FRES values indicated that the generated content generally corresponded to a difficult reading level.</p> Conclusion <p>The evaluated AI chatbot systems demonstrated largely comparable performance in responding to laminate veneer–related inquiries. While overall content quality was high, limitations in reliability and elevated readability levels indicate that AI-generated responses should be interpreted cautiously and used as supportive tools rather than substitutes for professional clinical guidance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of AI chatbots in responding to patients’ frequently asked questions about laminate veneer restorations

  • Bike Altan,
  • Şevki Çinar

摘要

Background

Although AI (artificial intelligence) chatbot systems are increasingly used to obtain dental information, limited evidence exists regarding the reliability and readability of their responses in laminate veneer–related patient inquiries. The aim of the study is to assess the reliability, quality, and readability of the responses given by 3 different artificial intelligence chatbot systems to frequently asked questions about laminate veneer restorations by patients.

Methods

Twenty-five frequently asked questions about laminate veneer restorations submitted to ChatGPT 5.2, Gemini 3, and DeepSeek-V3.2. AI chatbots to generate answers. The quality of these answers was assessed by two different experts using GQS (Global Quality Scale), reliability using modified DISCERN, and readability using FRES (Flesch Reading Ease Score) and FKGL (Flesch–Kincaid Grade Level). Normality was assessed with the Shapiro–Wilk test. Depending on data distribution, repeated-measures ANOVA or the Friedman test was used for comparisons among chatbots. Dunn’s test was applied for post hoc analyses. A p-value < 0.05 was considered statistically significant.

Results

No statistically significant differences were found among chatbot systems in terms of DISCERN, FRES, or FKGL scores (P > 0.05). Interobserver agreement for DISCERN was generally poor, with only Gemini demonstrating statistically significant agreement (ICC = 0.393; P = 0.023). Although a significant overall difference was observed in GQS scores (P = 0.015), no significant pairwise differences were detected. Readability analysis revealed that all chatbot responses exceeded the recommended 6th–8th grade reading level for patient education materials, with mean FKGL values corresponding to approximately 1–2 grade levels above the recommended upper limit. In addition, FRES values indicated that the generated content generally corresponded to a difficult reading level.

Conclusion

The evaluated AI chatbot systems demonstrated largely comparable performance in responding to laminate veneer–related inquiries. While overall content quality was high, limitations in reliability and elevated readability levels indicate that AI-generated responses should be interpreted cautiously and used as supportive tools rather than substitutes for professional clinical guidance.