Evaluating Artificial Intelligence-Generated Patient Education Materials for Bariatric Surgery: Comparative Analysis of Response Quality, Reliability, and Readability Across ChatGPT and DeepSeek Models
摘要
Artificial intelligence (AI) models such as ChatGPT and DeepSeek have gained increasing attention for their potential to enhance patient education by delivering accessible and evidence-based health information. We designed the following study to evaluate the AI models—ChatGPT and DeepSeek—in generating patient education materials for bariatric surgery.
MethodsThirty commonly asked patient questions related to bariatric surgery were classified into four thematic domains: (1) surgical planning and technical considerations, (2) preoperative assessment and optimization, (3) postoperative care and complication management, and (4) long-term follow-up and disease management. Responses generated by ChatGPT and DeepSeek were evaluated using three key metrics: (1) response quality, assessed by the Global Quality Score, rated on a 5-point scale from 1 (poor) to 5 (excellent); (2) reliability, measured using modified DISCERN criteria, which assess adherence to clinical guidelines and evidence-based standards, with scores ranging from 5 (low) to 25 (high); and (3) readability, evaluated using two validated formulas: the Flesch-Kincaid Grade Level and the Flesch Reading Ease Score.
ResultsChatGPT significantly outperformed DeepSeek in response quality, with a median (IQR) Global Quality Score of 5.00 (4.00, 5.00) vs. 4.00 (4.00, 5.00) (P = 0.002). Higher reliability was also observed in ChatGPT, as reflected by mDISCERN scores across all four domains (median [IQR], 22.0 [21.0, 23.25] vs. 19.7 [19.0, 20.75]; P < 0.001). While no significant difference was found in the Flesch Reading Ease Score (mean [SD], 26.11 [12.84] vs. 20.87 [12.20]; P = 0.110), ChatGPT yielded significantly higher Flesch-Kincaid Grade Level Scores (meaning its text was more complex) (mean [SD], 16.40 [2.43] vs. 13.48 [2.35]; P < 0.001). Both models produced responses at a readability level corresponding to college education.
ConclusionsChatGPT provided higher-quality and more reliable responses, while DeepSeek’s answers were slightly easier to read. However, both models’ answers lacked attention to psychosocial and cultural aspects of patient care, highlighting the need for more empathetic, adaptive AI to support inclusive patient education.