Purpose <p>ChatGPT and other large language models (LLMs) have gained immense popularity since their commercial release in 2022, with applications in various sectors including health care. We sought to evaluate their deployment in anesthesiology and critical care in a systematic review. Our aim was to describe the integration of LLMs in the field by showcasing and categorizing their current applications, assessing their performance in patient care, and reviewing application-specific ethical and practical challenges in deployment.</p> Methods <p>Respecting Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines, we systematically searched through PubMed®, Embase, the Cochrane Central Register of Controlled Trials, and Web of Science®, from inception until 1 August 2024. We extracted all papers investigating LLMs in anesthesiology or critical care and reporting results. We segmented the literature into major themes and highlighted key findings and limitations.</p> Results <p>From 480 retrieved articles, we included 45 papers. The evaluated models (GPT-4, GPT-3.5, Google Bard [now Gemini], LLaMA, and others) showed diverse applications in four segments: intensive care unit, patient education, medical education, and perioperative care. Large language models, especially newer models, are promising in predicting clinical scores, navigating simple clinical scenarios, and managing preoperative anxiety. Their performance remains below the clinician level in predicting outcomes, solving complex clinical scenarios (i.e., airway management), board examinations, and generating patient-directed documents, although newer models performed better than older ones.</p> Conclusion <p>While LLMs are not yet equipped to fully assist physicians in anesthesiology and critical care, they have significant potential, and their capabilities are rapidly improving. Supervised use for select tasks can streamline patient care. Further trials are warranted as new versions of models become available.</p> Study registration <p>PROSPERO (<a href="https://www.crd.york.ac.uk/PROSPERO/view/CRD42024567380">CRD42024567380</a>); first submitted 22 July 2024.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The applications of ChatGPT and other large language models in anesthesiology and critical care: a systematic review

  • Nicolas Daccache,
  • Joe Zako,
  • Louis Morisson,
  • Pascal Laferrière-Langlois

摘要

Purpose

ChatGPT and other large language models (LLMs) have gained immense popularity since their commercial release in 2022, with applications in various sectors including health care. We sought to evaluate their deployment in anesthesiology and critical care in a systematic review. Our aim was to describe the integration of LLMs in the field by showcasing and categorizing their current applications, assessing their performance in patient care, and reviewing application-specific ethical and practical challenges in deployment.

Methods

Respecting Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines, we systematically searched through PubMed®, Embase, the Cochrane Central Register of Controlled Trials, and Web of Science®, from inception until 1 August 2024. We extracted all papers investigating LLMs in anesthesiology or critical care and reporting results. We segmented the literature into major themes and highlighted key findings and limitations.

Results

From 480 retrieved articles, we included 45 papers. The evaluated models (GPT-4, GPT-3.5, Google Bard [now Gemini], LLaMA, and others) showed diverse applications in four segments: intensive care unit, patient education, medical education, and perioperative care. Large language models, especially newer models, are promising in predicting clinical scores, navigating simple clinical scenarios, and managing preoperative anxiety. Their performance remains below the clinician level in predicting outcomes, solving complex clinical scenarios (i.e., airway management), board examinations, and generating patient-directed documents, although newer models performed better than older ones.

Conclusion

While LLMs are not yet equipped to fully assist physicians in anesthesiology and critical care, they have significant potential, and their capabilities are rapidly improving. Supervised use for select tasks can streamline patient care. Further trials are warranted as new versions of models become available.

Study registration

PROSPERO (CRD42024567380); first submitted 22 July 2024.