Readability and source transparency of AI-generated health information on human metapneumovirus: A comparative evaluation of five chatbots
摘要
This study aimed to evaluate the readability and citation practices of artificial intelligence (AI)-generated responses to questions about human metapneumovirus, a respiratory virus of growing public health concern.
Subject and methodsFive widely used AI chatbots—ChatGPT-4, Copilot, Gemini, Claude.ai, and Grok—were prompted with 14 standardized questions based on official guidelines from the World Health Organization, the Centers for Disease Control and Prevention, and the Australian National Health and Medical Research Council. Responses were anonymized and assessed using six established readability metrics: Flesch–Kincaid Reading Ease and Flesch–Kincaid Grade Level, Gunning Fog Index, SMOG (Simple Measure of Gobbledygook) Index, Coleman–Liau Index, and Automated Readability Index. Scores were compared to standards recommended by the American Medical Association and the National Institutes of Health. Citation frequency and credibility were also analyzed.
ResultsAmong 70 chatbot responses, only one met the recommended readability level. Median readability scores ranged from grade 10.4 to 16.0, indicating high complexity. One chatbot generated the most readable content, while another scored lowest. Only two chatbots included source citations. One cited 68 reliable sources, primarily from health organizations and academic institutions, while the other referenced 31 sources of varying quality.
ConclusionAI-generated health content often exceeds recommended readability thresholds and lacks consistent citation practices. These issues may hinder understanding and trust. Improving default readability settings and integrating real-time citation features could enhance the accessibility and credibility of chatbot-based health communication.