Objective <p>This study aimed to comparatively evaluate the scientific approach, accuracy, and practical applicability of different versions of ChatGPT in generating training programs.</p> Method <p>Adopting a mixed-methods design, the study employed seven distinct instruction sets, each developed with input from seven experts possessing a minimum of 10 years of professional experience (Certified Strength and Conditioning Specialists (CSCS), academicians holding PhDs in Sports Sciences with research focus on exercise physiology and training methodology) in their respective fields. Using these instruction sets, three versions of ChatGPT (ChatGPT-3.5, ChatGPT-4o, and ChatGPT-4.1) were tasked with generating 12-week resistance training programs for a hypothetical, healthy young adult male with a moderate training background. Using a rubric scoring scale, the generated programs were systematically evaluated and scored separately based on the following criteria: compliance with the initial program request, inclusion of literature references, adherence to exercise variety and progressive loading principles, individualization of progression and justification, program modifications, inclusion of warm-up and cool-down components, injury risk considerations, presence of incorrect recommendations and practical applicability, and accessibility.To determine whether differences in mean scores were statistically significant, the Friedman non-parametric test was applied. When significant differences were identified, pairwise comparisons were conducted using the Wilcoxon signed-rank test to determine which groups accounted for these differences. In addition, qualitative data analysis was performed to explore expert evaluations in depth, employing both content analysis and descriptive analysis techniques.</p> Results <p>Statistically significant differences were identified among the ChatGPT versions examined in this study: between ChatGPT-4o and ChatGPT-3.5 (<i>p</i> = .018), ChatGPT-4.1 and ChatGPT-3.5 (<i>p</i> = .018), and ChatGPT-4.1 and ChatGPT-4o (<i>p</i> = .018). Expert content analysis further indicated that ChatGPT-4.1 produced responses that were more detailed, internally consistent, and better supported by scientific literature compared to the other versions.</p> Conclusion <p>Although the ChatGPT versions examined in this study exhibited certain limitations, they demonstrated the potential to deliver structured exercise programs aligned with established training principles and relevant scientific literature. Nonetheless, to ensure that AI-assisted training plans provide safe, evidence-based, and individualized content, the involvement of qualified human expertise remains essential.</p> Trial registration <p>No official trial registration number was assigned.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative evaluation of ChatGPT versions in training program design: scientific approach, accuracy, and practical applicability

  • Ayça Genç,
  • Gökhan Recep Aydın,
  • Murat Kasap,
  • Ali Özkan,
  • Alpamys Rakhymzhanov,
  • Hasan Hüseyin Gürcan,
  • Sevim Güllü,
  • Bilal Demirhan

摘要

Objective

This study aimed to comparatively evaluate the scientific approach, accuracy, and practical applicability of different versions of ChatGPT in generating training programs.

Method

Adopting a mixed-methods design, the study employed seven distinct instruction sets, each developed with input from seven experts possessing a minimum of 10 years of professional experience (Certified Strength and Conditioning Specialists (CSCS), academicians holding PhDs in Sports Sciences with research focus on exercise physiology and training methodology) in their respective fields. Using these instruction sets, three versions of ChatGPT (ChatGPT-3.5, ChatGPT-4o, and ChatGPT-4.1) were tasked with generating 12-week resistance training programs for a hypothetical, healthy young adult male with a moderate training background. Using a rubric scoring scale, the generated programs were systematically evaluated and scored separately based on the following criteria: compliance with the initial program request, inclusion of literature references, adherence to exercise variety and progressive loading principles, individualization of progression and justification, program modifications, inclusion of warm-up and cool-down components, injury risk considerations, presence of incorrect recommendations and practical applicability, and accessibility.To determine whether differences in mean scores were statistically significant, the Friedman non-parametric test was applied. When significant differences were identified, pairwise comparisons were conducted using the Wilcoxon signed-rank test to determine which groups accounted for these differences. In addition, qualitative data analysis was performed to explore expert evaluations in depth, employing both content analysis and descriptive analysis techniques.

Results

Statistically significant differences were identified among the ChatGPT versions examined in this study: between ChatGPT-4o and ChatGPT-3.5 (p = .018), ChatGPT-4.1 and ChatGPT-3.5 (p = .018), and ChatGPT-4.1 and ChatGPT-4o (p = .018). Expert content analysis further indicated that ChatGPT-4.1 produced responses that were more detailed, internally consistent, and better supported by scientific literature compared to the other versions.

Conclusion

Although the ChatGPT versions examined in this study exhibited certain limitations, they demonstrated the potential to deliver structured exercise programs aligned with established training principles and relevant scientific literature. Nonetheless, to ensure that AI-assisted training plans provide safe, evidence-based, and individualized content, the involvement of qualified human expertise remains essential.

Trial registration

No official trial registration number was assigned.