Objective <p>This study evaluated the reliability, quality, utility, and readability of ChatGPT-4o’s responses to frequently asked patient questions about diabetes technologies, addressing the potential of AI tools in patient education.</p> Methods <p>Twenty open-ended questions were submitted to ChatGPT-4o in both English and Turkish. Three endocrinologists independently rated responses using a modified DISCERN (mDISCERN) scale, Global Quality Scale (GQS), a Likert-based utility scale, and standard readability metrics (FKGL, GFI, FRE). Five responses were further analyzed before and after prompt modification (“explain in simpler terms”) to assess readability improvements.</p> Results <p>Responses in English demonstrated “fair” reliability (mean mDISCERN: 28.4 ± 1.6), with 75% rated as “high quality” (median GQS: 4) and “moderately useful” (median utility score: 3). Readability was limited (mean FKGL: 10.5, GFI: 11.1; FRE: 37.8), requiring college-level literacy. Prompting simplification improved readability significantly (mean FKGL: 4.9; FRE: 79.0), without altering content. Strong negative correlations were found between readability complexity and perceived quality (FKGL vs. GQS: <i>r</i> = − 0.449, <i>p</i> = 0.047). Turkish responses were less detailed and sometimes outdated, highlighting linguistic asymmetry.</p> Conclusion <p>ChatGPT-4o responses offer moderate-to-high reliability and quality but limited readability. Prompt engineering can enhance accessibility, and English responses are more robust than Turkish ones. ChatGPT-4o aligns more with the educational (DSME) than behavioral support (DSMS) components of diabetes care. While useful as a complementary educational tool, clinical oversight remains essential to ensure safety and contextual accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChatGPT-4o as a digital health tool for diabetes technology education: insights on reliability, quality, and readability

  • Selin Tekin,
  • Seda Hanife Oguz,
  • Selcuk Dagdelen

摘要

Objective

This study evaluated the reliability, quality, utility, and readability of ChatGPT-4o’s responses to frequently asked patient questions about diabetes technologies, addressing the potential of AI tools in patient education.

Methods

Twenty open-ended questions were submitted to ChatGPT-4o in both English and Turkish. Three endocrinologists independently rated responses using a modified DISCERN (mDISCERN) scale, Global Quality Scale (GQS), a Likert-based utility scale, and standard readability metrics (FKGL, GFI, FRE). Five responses were further analyzed before and after prompt modification (“explain in simpler terms”) to assess readability improvements.

Results

Responses in English demonstrated “fair” reliability (mean mDISCERN: 28.4 ± 1.6), with 75% rated as “high quality” (median GQS: 4) and “moderately useful” (median utility score: 3). Readability was limited (mean FKGL: 10.5, GFI: 11.1; FRE: 37.8), requiring college-level literacy. Prompting simplification improved readability significantly (mean FKGL: 4.9; FRE: 79.0), without altering content. Strong negative correlations were found between readability complexity and perceived quality (FKGL vs. GQS: r = − 0.449, p = 0.047). Turkish responses were less detailed and sometimes outdated, highlighting linguistic asymmetry.

Conclusion

ChatGPT-4o responses offer moderate-to-high reliability and quality but limited readability. Prompt engineering can enhance accessibility, and English responses are more robust than Turkish ones. ChatGPT-4o aligns more with the educational (DSME) than behavioral support (DSMS) components of diabetes care. While useful as a complementary educational tool, clinical oversight remains essential to ensure safety and contextual accuracy.