Background <p>The integration of artificial intelligence (AI) into medical practice opens up new frontiers for decision support, especially in intricate surgical procedures like one-anastomosis gastric bypass (OAGB). This study was designed to showcase the potential and performance of three AI models—ChatGPT-4.0, ChatGPT-Omni, and Gemini AI—in tackling complex clinical queries related to OAGB, thereby paving the way for a more efficient and effective surgical practice.</p> Methods <p>The study utilized a comprehensive query evaluation methodology comprising 180 questions for ChatGPT-4.0, ChatGPT-Omni, and Gemini AI models, equally divided among true/false, multiple-choice, open-ended, and case-scenario queries. These questions covered various aspects of OAGB surgery, including preoperative assessment, surgical technique, management of complications, and long-term outcomes.</p> Results <p>ChatGPT-Omni showed higher accuracy rates than Gemini AI and ChatGPT-4.0 in most question formats and difficulty levels (<i>p</i> &lt; 0.0001). However, the performance gap varied depending on the complexity and type of the queries. In true–false and multiple-choice formats, ChatGPT-Omni excelled, particularly in complex scenarios (<i>p</i> = 0.017). With a mean of 5.62 on a six-point scale, ChatGPT-Omni demonstrated exceptional capability in providing accurate and comprehensive answers to both open-ended and case scenarios. ChatGPT-Omni demonstrated the highest performance metrics, including precision (0.947), recall (0.857), and F1-score (0.9), although these values were dependent on the specific query format and type.</p> Conclusions <p>While ChatGPT-Omni demonstrated superior accuracy in many clinical queries related to OAGB, especially in simpler decision-making scenarios, it is crucial to underscore the need for additional validation in complex clinical settings. This cautionary note serves as a reminder of the current limitations of AI in surgery and the importance of ongoing research and validation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Performance of Artificial Intelligence in One Anastomosis Gastric Bypass Surgery: Comparative Efficacy of ChatGPT-4.0, ChatGPT-Omni, and Gemini AI

  • Erkan Aksoy

摘要

Background

The integration of artificial intelligence (AI) into medical practice opens up new frontiers for decision support, especially in intricate surgical procedures like one-anastomosis gastric bypass (OAGB). This study was designed to showcase the potential and performance of three AI models—ChatGPT-4.0, ChatGPT-Omni, and Gemini AI—in tackling complex clinical queries related to OAGB, thereby paving the way for a more efficient and effective surgical practice.

Methods

The study utilized a comprehensive query evaluation methodology comprising 180 questions for ChatGPT-4.0, ChatGPT-Omni, and Gemini AI models, equally divided among true/false, multiple-choice, open-ended, and case-scenario queries. These questions covered various aspects of OAGB surgery, including preoperative assessment, surgical technique, management of complications, and long-term outcomes.

Results

ChatGPT-Omni showed higher accuracy rates than Gemini AI and ChatGPT-4.0 in most question formats and difficulty levels (p < 0.0001). However, the performance gap varied depending on the complexity and type of the queries. In true–false and multiple-choice formats, ChatGPT-Omni excelled, particularly in complex scenarios (p = 0.017). With a mean of 5.62 on a six-point scale, ChatGPT-Omni demonstrated exceptional capability in providing accurate and comprehensive answers to both open-ended and case scenarios. ChatGPT-Omni demonstrated the highest performance metrics, including precision (0.947), recall (0.857), and F1-score (0.9), although these values were dependent on the specific query format and type.

Conclusions

While ChatGPT-Omni demonstrated superior accuracy in many clinical queries related to OAGB, especially in simpler decision-making scenarios, it is crucial to underscore the need for additional validation in complex clinical settings. This cautionary note serves as a reminder of the current limitations of AI in surgery and the importance of ongoing research and validation.