3D Multimodal Feature Model Based on Visual Prompt Tuning for Cerebral Stroke Classification
摘要
Stroke is a cerebral dysfunction caused by acute cerebrovascular events, and the critical window period for thrombolytic therapy is within 4.5 h. Traditional machine learning methods, such as SVM, random forest, and KNN, struggle to handle high-dimensional, complex data. Deep learning methods, such as CNNs and RNNs, are good at working with this kind of data, but require large labeled datasets and are prone to class imbalance problems. In this paper, we propose a 3D multimodal feature fusion model, which combines 3D structural images with point cloud data and Diffusion Weighted Imaging (DWI) to enhance the classification capability. The model adopts 3D ConvNeXt2List of 3D convolution, 3D GRN for global information modeling, and 3D Bi-LSTM for global integration. Tip tuning is used to reduce reliance on large data sets and to mitigate overfitting in small sample scenarios. The residual network structure is used to enhance the point cloud network. Experimental results show the method significantly enhances stroke onset time prediction accuracy, validating its effectiveness and potential application.