Large Language Models (LLMs) are increasingly utilized in tasks such as code generation for medical image analysis. However, their specific influence on study outcomes remains under-explored. In this study, we provide a comprehensive evaluation of various open- and closed-source LLMs, comparing their performance across medical imaging datasets. Each LLM was tasked with generating code for a U-Net-based baseline for a semantic segmentation task, guided by a tailored prompt.We evaluated each LLM’s generated model performance using the Dice coefficient and recorded all interactions with the LLM. Significant variations in baseline performance were observed among the LLMs, with differences of up to Δ 85.49% for the Bolus, 86.33% for the BAGLS, and 87.32% for the Brain Tumor test dataset. Additionally, we identified LLMs with minimal coding errors (best-performing LLMs: GPT o1 Preview and Claude 3.5 Sonnet with zero errors upon initial code execution; least-performing: Gemini 1.5 Pro and LlAMA 3.1 405B with 15 and 11 errors, respectively). In summary, careful selection of LLMs can significantly enhance medical image analysis code generation and establish reliable baselines for further algorithmic development.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-driven Baselines for Medical Image Segmentation

  • Jasmin Arjomandi,
  • Luisa Neubig,
  • Andreas M. Kist

摘要

Large Language Models (LLMs) are increasingly utilized in tasks such as code generation for medical image analysis. However, their specific influence on study outcomes remains under-explored. In this study, we provide a comprehensive evaluation of various open- and closed-source LLMs, comparing their performance across medical imaging datasets. Each LLM was tasked with generating code for a U-Net-based baseline for a semantic segmentation task, guided by a tailored prompt.We evaluated each LLM’s generated model performance using the Dice coefficient and recorded all interactions with the LLM. Significant variations in baseline performance were observed among the LLMs, with differences of up to Δ 85.49% for the Bolus, 86.33% for the BAGLS, and 87.32% for the Brain Tumor test dataset. Additionally, we identified LLMs with minimal coding errors (best-performing LLMs: GPT o1 Preview and Claude 3.5 Sonnet with zero errors upon initial code execution; least-performing: Gemini 1.5 Pro and LlAMA 3.1 405B with 15 and 11 errors, respectively). In summary, careful selection of LLMs can significantly enhance medical image analysis code generation and establish reliable baselines for further algorithmic development.