错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-shot performance of a general-purpose vision-language model for pediatric appendicitis diagnosis

  • Ceren Altintas Mese,
  • Ismail Mese,
  • Tugba Akinci D’Antonoli

摘要

Background

General-purpose vision-language models can analyze medical images without task-specific training, but their value for pediatric abdominal ultrasound is unknown.

Objective

To investigate the zero-shot performance of a general-purpose vision-language model to diagnose pediatric appendicitis using multimodal inputs, including images, report text, and clinical information.

Materials and methods

In this retrospective study, diagnostic capabilities of Llama 4 Maverick were evaluated on the Regensburg Pediatric Appendicitis Dataset. Five experiments were conducted using the following inputs as well as input combinations: images only, ultrasound report text only, images and ultrasound text, images and clinical data, entire multimodal input. Reference standard of appendicitis was defined based on histopathology in patients who underwent surgical resection and on clinical follow-up in patients managed conservatively. Performance was assessed at the patient level using accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1-score. Area under the receiver operating characteristic curve (AUROC) was evaluated secondarily as a measure of discrimination.

Results

Of 782 patients, 294 met inclusion criteria; based on data availability, four experiments are conducted using 293 patients and in the fifth experiment, 284. Images-only experiment showed very low specificity (13.5%) and had the poorest discrimination (AUROC, 0.567). Ultrasound report text improved classification performance, with specificity increasing to 63.5% while maintaining high sensitivity; discrimination was also strong (AUROC, 0.883). Adding images to ultrasound text or combining all available inputs did not yield further meaningful improvement. Differences between images-only experiment and text-containing experiments were statistically significant (P≤0.002).

Conclusion

Zero-shot visual interpretation of pediatric abdominal ultrasound by a general-purpose vision-language model is inadequate for safe appendicitis diagnosis. Performance is driven primarily by a structured sonographic imaging report, supporting a role for such models as text-based decision support tools rather than autonomous ultrasound image interpreters.

Graphical Abstract