A proof-of-concept study of surgical-VLM for surgical support in robotic surgery: contextual benchmarking against ChatGPT-5
摘要
Artificial intelligence (AI) has rapidly advanced in surgical applications. However, existing single-modality AI models relying solely on image input lack the ability to integrate anatomical understanding with clinical reasoning, which is essential for safe and actual surgical decision-making. We constructed an AI model with multimodal training combining visual and linguistic data and named Surgical Vision-Language Model (Surgical-VLM) for real-time surgical support.
MethodsWe analyzed surgical videos from 50 cases of robotic distal gastrectomy and extracted 50 still images per case, generating 10,000 vision–question–answer (VQA) pairs. The model was fine-tuned using the Large Language and Vision Assistant (LLaVA) framework. Model performance was assessed using the newly developed Surgical-VLM Bench, which evaluates appropriateness of expression, anatomical accuracy, and clinical usefulness on a 5-point scale, and the Bidirectional Encoder Representations from Transformers (BERT) score. The results were compared with those of ChatGPT-5.
ResultsSurgical-VLM achieved a mean total benchmark score of 11.33 ± 0.697 versus 10.72 ± 0.536 for ChatGPT-5. Surgical-VLM showed numerically higher mean scores in anatomical accuracy and clinical usefulness. The BERTScore was 0.768 ± 0.008 for Surgical-VLM and 0.804 ± 0.013 for ChatGPT-5.
ConclusionsThis proof-of-concept study demonstrates the feasibility of domain adaptation for a Surgical-VLM and proposes a clinically grounded benchmark for structured evaluation. In this pilot setting, the prototype generated context-aware responses and showed domain-dependent differences compared with a general-purpose multimodal model. Further validation with larger independent test sets, expanded VQA items, and external evaluators is required before any clinical use.