<p>Colorectal cancer (CRC) necessitates effective patient education, yet patients increasingly utilize Artificial Intelligence (AI) large language models (LLMs) for health information, raising concerns about quality and accessibility. This study evaluates the suitability of leading LLMs for generating CRC patient education materials. Eleven standardized prompts covering CRC topics from diagnosis to prognosis were posed to each LLM in separate sessions. Readability was measured using Flesch-Kincaid Reading Ease (FRES) and Grade Level (FKGL). Quality was assessed using the DISCERN instrument (16 criteria, max score 80) by two independent reviewers. Readability consistently exceeded recommended levels for patient materials, with FKGL scores typically corresponding to US grades 7–10. DISCERN scores indicated ‘fair’ quality overall (range 37–58,), with ERNIE X1 demonstrating the most consistent performance (48–58) among the tested models. Gemini-2.5Pro generated significantly longer and more complex responses (higher average words per sentence) with higher quality variability (37–57). Critically, none of the four models provided citations or information sources. Current LLMs can generate factually accurate basic CRC information but exhibit significant limitations for patient education. Outputs consistently fail to meet recommended readability standards, potentially hindering comprehension for many patients. While demonstrating ‘fair’ quality, the universal lack of source citations severely compromises transparency and trustworthiness. Substantial improvements, particularly in simplifying language and providing verifiable sources, are essential before these AI tools can be reliably and safely used as standalone resources for CRC patient education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing Patient Education Materials for Colorectal Cancer Generated by Four Large Language Models: Readability, Quality, and Transparency Challenges

  • Ming-yi Yuan,
  • Wen-qi Hong,
  • Ren-hao Hu,
  • Xiao-hua Jiang,
  • Shun Zhang

摘要

Colorectal cancer (CRC) necessitates effective patient education, yet patients increasingly utilize Artificial Intelligence (AI) large language models (LLMs) for health information, raising concerns about quality and accessibility. This study evaluates the suitability of leading LLMs for generating CRC patient education materials. Eleven standardized prompts covering CRC topics from diagnosis to prognosis were posed to each LLM in separate sessions. Readability was measured using Flesch-Kincaid Reading Ease (FRES) and Grade Level (FKGL). Quality was assessed using the DISCERN instrument (16 criteria, max score 80) by two independent reviewers. Readability consistently exceeded recommended levels for patient materials, with FKGL scores typically corresponding to US grades 7–10. DISCERN scores indicated ‘fair’ quality overall (range 37–58,), with ERNIE X1 demonstrating the most consistent performance (48–58) among the tested models. Gemini-2.5Pro generated significantly longer and more complex responses (higher average words per sentence) with higher quality variability (37–57). Critically, none of the four models provided citations or information sources. Current LLMs can generate factually accurate basic CRC information but exhibit significant limitations for patient education. Outputs consistently fail to meet recommended readability standards, potentially hindering comprehension for many patients. While demonstrating ‘fair’ quality, the universal lack of source citations severely compromises transparency and trustworthiness. Substantial improvements, particularly in simplifying language and providing verifiable sources, are essential before these AI tools can be reliably and safely used as standalone resources for CRC patient education.