This study evaluates the effectiveness of AI-powered conversational agents (ChatGPT4mini, ChatGPT4o, Gemini, and Copilot) in delivering patient information on appropriate wheelchairs. Using the Ensuring Quality Information for Patients (EQIP) tool, the AI models were tested on 35 standardized questions, with two independent experts validating their responses. Readability was assessed using 11 Grammarly metrics across word count, vocabulary, and overall readability. Test-retest and inter-rater reliability were calculated using ICC while prompting interactions were analyzed using Kruskal-Wallis tests. Results showed an average quality score of 22.1/35, readability at 90.3%, and reliability at 0.98, with ChatGPT4o delivering the best performance across quality (25.8 ± 1.5), readability (94 ± 0%), and reliability (>0.99). Gemini scored the lowest in quality (14.3 ± 7.8), while Copilot had the lowest readability (87 ± 6.4%) and reliability (0.95 ± 0.07). Advanced prompts did not significantly improve outcomes. Despite promising results, the study highlights limitations like model constraints and training data gaps, emphasizing the need for further refinement and prompt engineering optimization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessment of the Quality and Readability of AI-Generated Wheelchair Information

  • Alberto Isaac Perez Sanpablo,
  • Alicia Meneses Peñaloza

摘要

This study evaluates the effectiveness of AI-powered conversational agents (ChatGPT4mini, ChatGPT4o, Gemini, and Copilot) in delivering patient information on appropriate wheelchairs. Using the Ensuring Quality Information for Patients (EQIP) tool, the AI models were tested on 35 standardized questions, with two independent experts validating their responses. Readability was assessed using 11 Grammarly metrics across word count, vocabulary, and overall readability. Test-retest and inter-rater reliability were calculated using ICC while prompting interactions were analyzed using Kruskal-Wallis tests. Results showed an average quality score of 22.1/35, readability at 90.3%, and reliability at 0.98, with ChatGPT4o delivering the best performance across quality (25.8 ± 1.5), readability (94 ± 0%), and reliability (>0.99). Gemini scored the lowest in quality (14.3 ± 7.8), while Copilot had the lowest readability (87 ± 6.4%) and reliability (0.95 ± 0.07). Advanced prompts did not significantly improve outcomes. Despite promising results, the study highlights limitations like model constraints and training data gaps, emphasizing the need for further refinement and prompt engineering optimization.