Background <p>Up to 30% of imaging examinations are deemed inappropriate, leading to unnecessary radiation exposure, emergency department (ED) overcrowding, and rising healthcare costs. While clinical decision support systems (CDSSs) such as ESR iGuide aim to improve imaging appropriateness, their clinical impact and user adoption remain limited. Large language models (LLMs) may offer a faster, more accessible alternative for radiologic decision-making.</p> Purpose <p>To evaluate the performance of OpenAccessGPT in reducing inappropriate radiological examinations requested from an ED by comparing its recommendations with ESR iGuide and assessing appropriateness and time efficiency.</p> Materials and methods <p>This retrospective, single-center, semi-quantitative cohort study was conducted at a tertiary hospital in Switzerland. A total of 201 consecutive ED radiology requests (March 2023–April 2024) met inclusion criteria (50 CT, 50 US, 50 X-rays, 51 MRI). Four hypotheses were tested: (1) OpenAccessGPT’s ability to reduce imaging requests; (2) concordance with ESR iGuide; (3) time efficiency; and (4) the rate of inappropriate examinations in clinical practice when confronted to ESR iGuide. Statistical analyses included proportion tests and paired t-tests (<i>p</i> &lt; .05).</p> Results <p>Mean patient age was 55.3 years (SD = 20.9); 45% were female. Eleven cases were excluded due to unavailable ESR iGuide scenarios. OpenAccessGPT advised against imaging in 6% of cases (12/201), with 33.3% agreement with ESR iGuide (5/12). Overall concordance was 78.9% (150/190), below the 90% threshold (χ² = 25.79, <i>p</i> &lt; .001). GPT was significantly faster (21.6&#xa0;s vs. 101&#xa0;s; Δ = 79.6&#xa0;s; t[189] = 12.8; <i>p</i> &lt; .001). Institutional practice aligned with ESR iGuide in 82.6% of cases (χ² = 11.46, <i>p</i> &lt; .001).</p> Conclusion <p>OpenAccessGPT showed only limited ability to reduce inappropriate imaging. Despite substantial faster response times, its clinical reliability remained insufficient for standalone use.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI as a clinical decision support system: accuracy of an open GPT-model in emergency practice

  • Jérémie Arthaud,
  • Harriet C. Thoeny,
  • Vincent Grek,
  • Lucien Widmer

摘要

Background

Up to 30% of imaging examinations are deemed inappropriate, leading to unnecessary radiation exposure, emergency department (ED) overcrowding, and rising healthcare costs. While clinical decision support systems (CDSSs) such as ESR iGuide aim to improve imaging appropriateness, their clinical impact and user adoption remain limited. Large language models (LLMs) may offer a faster, more accessible alternative for radiologic decision-making.

Purpose

To evaluate the performance of OpenAccessGPT in reducing inappropriate radiological examinations requested from an ED by comparing its recommendations with ESR iGuide and assessing appropriateness and time efficiency.

Materials and methods

This retrospective, single-center, semi-quantitative cohort study was conducted at a tertiary hospital in Switzerland. A total of 201 consecutive ED radiology requests (March 2023–April 2024) met inclusion criteria (50 CT, 50 US, 50 X-rays, 51 MRI). Four hypotheses were tested: (1) OpenAccessGPT’s ability to reduce imaging requests; (2) concordance with ESR iGuide; (3) time efficiency; and (4) the rate of inappropriate examinations in clinical practice when confronted to ESR iGuide. Statistical analyses included proportion tests and paired t-tests (p < .05).

Results

Mean patient age was 55.3 years (SD = 20.9); 45% were female. Eleven cases were excluded due to unavailable ESR iGuide scenarios. OpenAccessGPT advised against imaging in 6% of cases (12/201), with 33.3% agreement with ESR iGuide (5/12). Overall concordance was 78.9% (150/190), below the 90% threshold (χ² = 25.79, p < .001). GPT was significantly faster (21.6 s vs. 101 s; Δ = 79.6 s; t[189] = 12.8; p < .001). Institutional practice aligned with ESR iGuide in 82.6% of cases (χ² = 11.46, p < .001).

Conclusion

OpenAccessGPT showed only limited ability to reduce inappropriate imaging. Despite substantial faster response times, its clinical reliability remained insufficient for standalone use.