Introduction and Hypothesis <p>The objective was to develop a retrieval-augmented ChatGPT model grounded in evidence-based patient education materials and compare its performance against the standard ChatGPT model in responding to common urogynecology patient questions in this pilot study.</p> Methods <p>We developed a retrieval-augmented ChatGPT-4.0 model that prioritized content from International Urogynecological Association patient information leaflets. Ten commonly asked patient questions were submitted to both the standard and retrieval-augmented models. Six board-certified urogynecologists evaluated responses using the validated Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, which assesses accuracy, clarity, relevance, completeness, usefulness, and sources. Total and domain-specific QAMAI scores were compared using the Wilcoxon signed-rank test, and a sensitivity analysis was performed, excluding the unblinded Source domain.</p> Results <p>The retrieval-augmented model demonstrated significantly higher total QAMAI scores than the standard model (median 22 [interquartile range, IQR, 19–25] vs 16 [IQR 13–18], <i>p</i> &lt; 0.01) and outperformed the standard model in all six domains.&#xa0;In the sensitivity analysis, the retrieval-augmented model maintained significantly higher performance (18 [IQR 16–20] vs 14.5 [IQR 11–17], <i>p</i> &lt; 0.01).&#xa0;Clinician raters preferred the retrieval-augmented model in 81% of responses.</p> Conclusions <p>Grounding AI tools in vetted patient education materials significantly improved the quality of ChatGPT-generated responses in urogynecology. Retrieval-augmented models offer a promising approach to enhance patient education and promote patient-centered care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development and Evaluation of an Augmented Artificial Intelligence Model for Urogynecology Queries

  • Madeline K. Moureau,
  • Berkley Davis,
  • Christopher X. Hong

摘要

Introduction and Hypothesis

The objective was to develop a retrieval-augmented ChatGPT model grounded in evidence-based patient education materials and compare its performance against the standard ChatGPT model in responding to common urogynecology patient questions in this pilot study.

Methods

We developed a retrieval-augmented ChatGPT-4.0 model that prioritized content from International Urogynecological Association patient information leaflets. Ten commonly asked patient questions were submitted to both the standard and retrieval-augmented models. Six board-certified urogynecologists evaluated responses using the validated Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, which assesses accuracy, clarity, relevance, completeness, usefulness, and sources. Total and domain-specific QAMAI scores were compared using the Wilcoxon signed-rank test, and a sensitivity analysis was performed, excluding the unblinded Source domain.

Results

The retrieval-augmented model demonstrated significantly higher total QAMAI scores than the standard model (median 22 [interquartile range, IQR, 19–25] vs 16 [IQR 13–18], p < 0.01) and outperformed the standard model in all six domains. In the sensitivity analysis, the retrieval-augmented model maintained significantly higher performance (18 [IQR 16–20] vs 14.5 [IQR 11–17], p < 0.01). Clinician raters preferred the retrieval-augmented model in 81% of responses.

Conclusions

Grounding AI tools in vetted patient education materials significantly improved the quality of ChatGPT-generated responses in urogynecology. Retrieval-augmented models offer a promising approach to enhance patient education and promote patient-centered care.