<p>Artificial intelligence (AI) tools like ChatGPT-4o are increasingly utilized in prenatal care. However, their reliability and clinical applicability for healthcare providers in first-trimester screening remain unclear. This study aimed to evaluate the reliability, readability, and clinical utility of ChatGPT-4o’s responses to support clinicians in counseling regarding combined first-trimester screening and non-invasive prenatal testing (NIPT). Fifteen risk-stratified clinical scenarios were used to prompt ChatGPT-4o. Fourteen perinatologists rated the responses using mDISCERN and Global Quality Scale (GQS). Readability was assessed via five indices. Inter-rater agreement and internal consistency were evaluated using ICC and Cronbach’s alpha. AI responses showed high inter-rater reliability (ICC = 0.998) and internal consistency (α = 0.975). GQS and mDISCERN scores were highest in high-risk scenarios. Readability did not significantly differ across risk levels, nor correlate with quality scores. ChatGPT-4o demonstrates potential as a clinical decision-support and counseling tool for clinicians involved in prenatal screening, particularly in high-risk scenarios. Further refinement is needed for consistent performance across risk levels.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the reliability and clinical utility of artificial intelligence in first trimester prenatal screening and noninvasive prenatal testing

  • İbrahim Taşkum,
  • Selcan Sınacı,
  • Seyhun Sucu,
  • Fatma Didem Yücel Yetişkin

摘要

Artificial intelligence (AI) tools like ChatGPT-4o are increasingly utilized in prenatal care. However, their reliability and clinical applicability for healthcare providers in first-trimester screening remain unclear. This study aimed to evaluate the reliability, readability, and clinical utility of ChatGPT-4o’s responses to support clinicians in counseling regarding combined first-trimester screening and non-invasive prenatal testing (NIPT). Fifteen risk-stratified clinical scenarios were used to prompt ChatGPT-4o. Fourteen perinatologists rated the responses using mDISCERN and Global Quality Scale (GQS). Readability was assessed via five indices. Inter-rater agreement and internal consistency were evaluated using ICC and Cronbach’s alpha. AI responses showed high inter-rater reliability (ICC = 0.998) and internal consistency (α = 0.975). GQS and mDISCERN scores were highest in high-risk scenarios. Readability did not significantly differ across risk levels, nor correlate with quality scores. ChatGPT-4o demonstrates potential as a clinical decision-support and counseling tool for clinicians involved in prenatal screening, particularly in high-risk scenarios. Further refinement is needed for consistent performance across risk levels.