Multimodal Breast Ultrasound Segmentation: Combining Visual and Clinical Data
摘要
Breast cancer remains one of the leading causes of cancer-related deaths among women worldwide, highlighting the critical need for early detection and accurate diagnosis. Breast ultrasound (BUS) imaging is one of the most essential methods for early diagnosis of breast cancer. In this research, we develop a novel hybrid U-shaped network for the automated segmentation of breast lesions in BUS images. Meanwhile, we employ a multimodal approach that combines visual features from ultrasound images with contextual textual information, improving the model’s understanding of tumor characteristics. We evaluate various configurations, including selecting clinical features such as tumor classification and BI-RADS scores. Our findings show that the proposed multimodal framework outperforms some transformer-based models on a public BUS image dataset, achieving significant advancements. This study underscores the effectiveness of multimodal learning in medical image analysis and highlights the potential of transformer-language models to improve diagnostic tools for early breast cancer detection.