<p>This study evaluates the reliability, completeness, and thematic consistency of responses generated by legacy versions of ChatGPT—3.5 (2023) and 4.0 (2024)—across five autism-related domains: Diagnosis, Prognosis, Prevalence, Evaluation, and Treatment. The study provides a historical snapshot of model performance while highlighting broader patterns of accuracy, limitations, and implications for ongoing use of large language models in autism-related contexts. Sixty-nine questions across five domains were presented to both ChatGPT-3.5 and ChatGPT-4.0. The responses were evaluated for accuracy, length of explanation, completeness, and thematic consistency. Comparative analyses incorporated both descriptive and inferential statistics, and thematic patterns were examined using cosine similarity. Both ChatGPT-3.5 and ChatGPT-4.0 exhibited high accuracy, particularly in structured domains such as Diagnosis and Treatment. ChatGPT-4.0 provided slightly richer descriptive detail in more complex areas, though this occasionally introduced thematic variability. Completeness scores were moderate across domains (ranging from ~ 0.20 to 0.50), reflecting that responses often captured some, but not all expected key points. A limited consistency checks with ChatGPT-5.0 (September 2025) demonstrated broad stability of conclusions, with no decreases in accuracy and only minor updates observed in prevalence-related answers. ChatGPT-3.5 and ChatGPT-4.0 show substantial potential as supplementary resources for autism-related information. Although these models consistently provide accurate content, their moderate completeness underscores the need for human oversight to ensure quality and comprehensiveness. Understanding their reliability, limitations, and evolving nature is essential as large language models continue to be explored for healthcare education and patient-facing information, though not as replacements for professional training or clinical care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An empirical evaluation of ChatGPT-3.5 and ChatGPT-4.0 for autism-related queries: validity, completeness, and consistency

  • Ali Naderi Malek,
  • Patricia Prelock,
  • Atefeh Jannesari,
  • Fatemeh Mehrpour

摘要

This study evaluates the reliability, completeness, and thematic consistency of responses generated by legacy versions of ChatGPT—3.5 (2023) and 4.0 (2024)—across five autism-related domains: Diagnosis, Prognosis, Prevalence, Evaluation, and Treatment. The study provides a historical snapshot of model performance while highlighting broader patterns of accuracy, limitations, and implications for ongoing use of large language models in autism-related contexts. Sixty-nine questions across five domains were presented to both ChatGPT-3.5 and ChatGPT-4.0. The responses were evaluated for accuracy, length of explanation, completeness, and thematic consistency. Comparative analyses incorporated both descriptive and inferential statistics, and thematic patterns were examined using cosine similarity. Both ChatGPT-3.5 and ChatGPT-4.0 exhibited high accuracy, particularly in structured domains such as Diagnosis and Treatment. ChatGPT-4.0 provided slightly richer descriptive detail in more complex areas, though this occasionally introduced thematic variability. Completeness scores were moderate across domains (ranging from ~ 0.20 to 0.50), reflecting that responses often captured some, but not all expected key points. A limited consistency checks with ChatGPT-5.0 (September 2025) demonstrated broad stability of conclusions, with no decreases in accuracy and only minor updates observed in prevalence-related answers. ChatGPT-3.5 and ChatGPT-4.0 show substantial potential as supplementary resources for autism-related information. Although these models consistently provide accurate content, their moderate completeness underscores the need for human oversight to ensure quality and comprehensiveness. Understanding their reliability, limitations, and evolving nature is essential as large language models continue to be explored for healthcare education and patient-facing information, though not as replacements for professional training or clinical care.