Exploring the accuracy and consistency of AI models in dyslexia: evaluating ChatGPT and gemini as supplementary resources for parents and caregivers
摘要
As artificial intelligence becomes more widespread, tools such as GPT-3.5, GPT-4, and Google Gemini are also increasingly used as additional resources to help parents understand dyslexia. However, the validity and precision of these instruments have not been studied well. This cross-sectional study assessed ChatGPT-3.5, ChatGPT-4, and Gemini on 107 dyslexia-related questions curated from institutional and social media sources, categorized into General Knowledge, Management, and Other domains. Accuracy, comprehensiveness, reproducibility, and reliability were evaluated using a 4-point accuracy scale and Chi-square analyses, with R software for statistical and visual outputs. GPT-4 demonstrated the best correct and complete responses in every category, performing particularly well in “Management” and “General Knowledge”. GPT-3.5 performed exceptionally well in challenging questions, albeit with less detail, while Google Gemini had a competitive but somewhat uneven performance in the “Other” category. All models were reproducible, and GPT-4 demonstrated 100% reproducibility in initial knowledge and management questions. For all models, the interrater agreement was more than 99%. These findings support integrating AI-driven assistants—particularly GPT-4—into parental support platforms, school and clinical portals, and self-help resources to deliver accurate, accessible guidance on dyslexia screening, intervention strategies, and accommodations for families managing dyslexia.
Graphical Abstract