错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A proof-of-concept study of surgical-VLM for surgical support in robotic surgery: contextual benchmarking against ChatGPT-5

  • Jumpei Ikeda,
  • Hirofumi Kawakubo,
  • Masashi Takeuchi,
  • Yosuke Morimoto,
  • Kazuaki Matsui,
  • Satoru Matsuda,
  • Yuko Kitagawa

摘要

Background

Artificial intelligence (AI) has rapidly advanced in surgical applications. However, existing single-modality AI models relying solely on image input lack the ability to integrate anatomical understanding with clinical reasoning, which is essential for safe and actual surgical decision-making. We constructed an AI model with multimodal training combining visual and linguistic data and named Surgical Vision-Language Model (Surgical-VLM) for real-time surgical support.

Methods

We analyzed surgical videos from 50 cases of robotic distal gastrectomy and extracted 50 still images per case, generating 10,000 vision–question–answer (VQA) pairs. The model was fine-tuned using the Large Language and Vision Assistant (LLaVA) framework. Model performance was assessed using the newly developed Surgical-VLM Bench, which evaluates appropriateness of expression, anatomical accuracy, and clinical usefulness on a 5-point scale, and the Bidirectional Encoder Representations from Transformers (BERT) score. The results were compared with those of ChatGPT-5.

Results

Surgical-VLM achieved a mean total benchmark score of 11.33 ± 0.697 versus 10.72 ± 0.536 for ChatGPT-5. Surgical-VLM showed numerically higher mean scores in anatomical accuracy and clinical usefulness. The BERTScore was 0.768 ± 0.008 for Surgical-VLM and 0.804 ± 0.013 for ChatGPT-5.

Conclusions

This proof-of-concept study demonstrates the feasibility of domain adaptation for a Surgical-VLM and proposes a clinically grounded benchmark for structured evaluation. In this pilot setting, the prototype generated context-aware responses and showed domain-dependent differences compared with a general-purpose multimodal model. Further validation with larger independent test sets, expanded VQA items, and external evaluators is required before any clinical use.