A Vision-Language Framework for Multimodal Retrieval in Glioma Histopathology
摘要
Gliomas remain among the most diagnostically challenging tumors of the central nervous system, owing to marked morphological heterogeneity and significant inter-observer variability in histopathological interpretation. Whole-slide images (WSIs) provide detailed morphologic information but are limited by their extreme resolution and susceptibility to inter-slide stain variability, complicating computational analysis. In a collaborative effort between our neuropathology and computational teams, we curated a novel high-quality dataset of 228 WSIs scanned at 40 \(\times \) magnification, from which 75,720 hematoxylin and eosin (H&E) patches were systematically extracted across six WHO-recognized glioma subtypes. Each patch was paired with an anonymized, neuropathologist authored diagnostic description, enabling clinically reliable image and text alignment. Building on this foundation, we present a novel vision–language contrastive framework that amalgamates a Vision Transformer (ViT-B/16) for morphological encoding with BioClinicalBERT for clinical text representation, jointly optimized in a shared embedding space via symmetric InfoNCE loss. The model achieved a Recall@10 of 55%, markedly outperforming the CLIP baseline, yet leaves scope for improvement in text-to-image and image-to-image retrieval tasks in neuropathology. Qualitative analyses further demonstrate the ability to capture fine-grained morphologic linguistic correspondences and establish a benchmark for multimodal retrieval in computational neuropathology.