ChatGPT-4o as a digital health tool for diabetes technology education: insights on reliability, quality, and readability
摘要
This study evaluated the reliability, quality, utility, and readability of ChatGPT-4o’s responses to frequently asked patient questions about diabetes technologies, addressing the potential of AI tools in patient education.
MethodsTwenty open-ended questions were submitted to ChatGPT-4o in both English and Turkish. Three endocrinologists independently rated responses using a modified DISCERN (mDISCERN) scale, Global Quality Scale (GQS), a Likert-based utility scale, and standard readability metrics (FKGL, GFI, FRE). Five responses were further analyzed before and after prompt modification (“explain in simpler terms”) to assess readability improvements.
ResultsResponses in English demonstrated “fair” reliability (mean mDISCERN: 28.4 ± 1.6), with 75% rated as “high quality” (median GQS: 4) and “moderately useful” (median utility score: 3). Readability was limited (mean FKGL: 10.5, GFI: 11.1; FRE: 37.8), requiring college-level literacy. Prompting simplification improved readability significantly (mean FKGL: 4.9; FRE: 79.0), without altering content. Strong negative correlations were found between readability complexity and perceived quality (FKGL vs. GQS: r = − 0.449, p = 0.047). Turkish responses were less detailed and sometimes outdated, highlighting linguistic asymmetry.
ConclusionChatGPT-4o responses offer moderate-to-high reliability and quality but limited readability. Prompt engineering can enhance accessibility, and English responses are more robust than Turkish ones. ChatGPT-4o aligns more with the educational (DSME) than behavioral support (DSMS) components of diabetes care. While useful as a complementary educational tool, clinical oversight remains essential to ensure safety and contextual accuracy.