A CNN-based liver ultrasound plane recognition system with multimodal large language model-generated interpretable feedback
摘要
To construct a hybrid artificial intelligence (AI) pipeline that combines a convolutional neural network (CNN) for image recognition with a multimodal large language model (MLLM) for textual interpretation, and to evaluate its application value in identifying liver ultrasound standard planes (LUSPs).
Materials and methodsThis multicenter study included 16,782 images in the standard dataset (training: 13508; testing: 3274), an internal validation set (3431 images), and an external validation set (5253 images). Using expert consensus as the reference standard, diagnostic performance was compared between students, junior radiologists, intermediate radiologists, senior radiologists, the AI model, and the same readers with AI assistance.
ResultsThe CNN component of the system recognized 13 LUSPs and nonstandard planes (N-SPs) with an accuracy of 88.8% in the test set. In the validation set, the area under the receiver operating curve, sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and kappa value of the AI model in identifying liver ultrasound planes were similar to those of senior radiologists (all p > 0.05). Except for the positive predictive value, the AI model outperformed students, junior radiologists, and intermediate radiologists in other diagnostic performance and these metrics also showed statistically significant improvement with hybrid AI system assistance under the conditions of this experimental setting (all p < 0.05). The model demonstrated robust generalizability across multiple centers, ultrasound devices, and diverse clinical scenarios.
ConclusionThis hybrid AI system, combining CNN-based image identification with MLLM-generated interpretable feedback, demonstrates potential as an effective auxiliary tool for image quality control and standardized training, warranting further prospective validation.