Background <p>The use of artificially intelligent Language Models (LLMs) like ChatGPT is increasing rapidly in medicine, but their accuracy, reasoning quality, and contextual safety remains an issue. This is especially relevant for forensic pathology, where balanced reasoning, contextual sensitivity, and precise communication are essential. We performed an in-depth assessment of ChatGPT’s capabilities and limitations for forensic pathology, which also improves our understanding of the risks and benefits of LLM use in medicine more broadly.</p> Methods <p>ChatGPT-4.5 Turbo’s performance was tested using a multifaceted mock exam, consisting of core elements of the forensic pathology fellowship exam of the Royal College of Pathologists of Australasia (RCPA). The mock exam included essay-style questions, image-based tasks, and case reporting. ChatGPT’s responses were blindly marked by experienced examiners according to standard RCPA criteria, assessing factual accuracy, reasoning structure, and communication quality.</p> Results <p>ChatGPT performed well on most essay-style knowledge questions, achieving higher scores on topics with well-established knowledge. Performance was however poor for tasks requiring complex reasoning, image interpretation, or the context-dependent analysis of autopsy findings. Importantly, ChatGPT’s output was always phrased fluently and persuasively, creating an impression of confidence that was independent of factual accuracy.</p> Conclusions <p>ChatGPT can reliably reproduce well-established forensic pathology knowledge. However, it lacks the capabilities needed for higher-level tasks and often generates unjustifiably confident and misleading output. Its use may be acceptable for low-risk administrative or educational purposes, provided output is carefully reviewed by qualified experts. Non-specialists should not rely on such tools for forensic pathology information. Continued evaluation is needed to ensure safe and responsible use.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChatGPT’s performance on a specialist forensic pathology examination: implications for forensic pathologists and non-specialists

  • Hans H. de Boer,
  • Gregory Young,
  • Heinrich Bouwer,
  • Karen J. Heath

摘要

Background

The use of artificially intelligent Language Models (LLMs) like ChatGPT is increasing rapidly in medicine, but their accuracy, reasoning quality, and contextual safety remains an issue. This is especially relevant for forensic pathology, where balanced reasoning, contextual sensitivity, and precise communication are essential. We performed an in-depth assessment of ChatGPT’s capabilities and limitations for forensic pathology, which also improves our understanding of the risks and benefits of LLM use in medicine more broadly.

Methods

ChatGPT-4.5 Turbo’s performance was tested using a multifaceted mock exam, consisting of core elements of the forensic pathology fellowship exam of the Royal College of Pathologists of Australasia (RCPA). The mock exam included essay-style questions, image-based tasks, and case reporting. ChatGPT’s responses were blindly marked by experienced examiners according to standard RCPA criteria, assessing factual accuracy, reasoning structure, and communication quality.

Results

ChatGPT performed well on most essay-style knowledge questions, achieving higher scores on topics with well-established knowledge. Performance was however poor for tasks requiring complex reasoning, image interpretation, or the context-dependent analysis of autopsy findings. Importantly, ChatGPT’s output was always phrased fluently and persuasively, creating an impression of confidence that was independent of factual accuracy.

Conclusions

ChatGPT can reliably reproduce well-established forensic pathology knowledge. However, it lacks the capabilities needed for higher-level tasks and often generates unjustifiably confident and misleading output. Its use may be acceptable for low-risk administrative or educational purposes, provided output is carefully reviewed by qualified experts. Non-specialists should not rely on such tools for forensic pathology information. Continued evaluation is needed to ensure safe and responsible use.