<p>Medical misinformation is a major public health concern. The public increasingly uses artificial intelligence (AI) tools for medical consultations. Therefore, concerns arise about their ability to detect and even correct subtle medical information that users may be embedding in users prompts. This study assessed the ability of different ChatGPT models in detecting and correcting such subtle misinformation.&#xa0;Fifty clinical plausible prompts with subtle medical misinformation were introduced separately to ChatGPT models 4o, 4.1-mini, and GPT-5. Prompts spanned Internal Medicine, Cardiology, Pediatrics, Ophthalmology, and Oncology. Responses were scored on a 3-point scale: 0: No correction; 1: Hedging or uncertainty; 3: cutting edge detection and correction.&#xa0;GPT-4o was the best performing model, surpassing GPT-5 by correctly identifying and correcting misinformation in 86% of the prompts compared to 74% for GPT-5. GPT-4.1-mini showed weaker performance, detecting dsmisinformation in only 52% of prompts, with complete failure in 34% and hedging in 14%. Specialty-specific analysis revealed that GPT-4o achieved higher detection rate in all tested specialties compared to GPT-4.1-mini and GPT-5. Only oncology showed comparable detection rates between GPT-4o and GPT-5.&#xa0;Although the performance of GPT-4o and GPT-5 in detecting subtle medical misinformation was promising, unexpectedly, GPT-4o surpassed GPT-5 in performance. Using underpowered variants such as GPT-4.1-mini, poses a public health threat. Reverse prompting offers a diagnostic lens and should be integrated into standard AI safety testing protocols.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial Intelligence’s Capacity to Detect Subtle Medical Misinformation: A Novel Reverse Prompting Approach

  • Mohamed Bendary,
  • Nouran Ramzy,
  • Amira Khater,
  • Mahmud Magdy Nasif,
  • Nora Atef

摘要

Medical misinformation is a major public health concern. The public increasingly uses artificial intelligence (AI) tools for medical consultations. Therefore, concerns arise about their ability to detect and even correct subtle medical information that users may be embedding in users prompts. This study assessed the ability of different ChatGPT models in detecting and correcting such subtle misinformation. Fifty clinical plausible prompts with subtle medical misinformation were introduced separately to ChatGPT models 4o, 4.1-mini, and GPT-5. Prompts spanned Internal Medicine, Cardiology, Pediatrics, Ophthalmology, and Oncology. Responses were scored on a 3-point scale: 0: No correction; 1: Hedging or uncertainty; 3: cutting edge detection and correction. GPT-4o was the best performing model, surpassing GPT-5 by correctly identifying and correcting misinformation in 86% of the prompts compared to 74% for GPT-5. GPT-4.1-mini showed weaker performance, detecting dsmisinformation in only 52% of prompts, with complete failure in 34% and hedging in 14%. Specialty-specific analysis revealed that GPT-4o achieved higher detection rate in all tested specialties compared to GPT-4.1-mini and GPT-5. Only oncology showed comparable detection rates between GPT-4o and GPT-5. Although the performance of GPT-4o and GPT-5 in detecting subtle medical misinformation was promising, unexpectedly, GPT-4o surpassed GPT-5 in performance. Using underpowered variants such as GPT-4.1-mini, poses a public health threat. Reverse prompting offers a diagnostic lens and should be integrated into standard AI safety testing protocols.