错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Vision-Language Framework for Multimodal Retrieval in Glioma Histopathology

  • Sagar Kumar Jha,
  • Abhishek Rana,
  • Jyotsna Singh,
  • Rohan Beriwal,
  • Sagnik Jana,
  • Madhavi Mathur,
  • Afza Akbar,
  • Vaishali Suri,
  • Tavpritesh Sethi

摘要

Gliomas remain among the most diagnostically challenging tumors of the central nervous system, owing to marked morphological heterogeneity and significant inter-observer variability in histopathological interpretation. Whole-slide images (WSIs) provide detailed morphologic information but are limited by their extreme resolution and susceptibility to inter-slide stain variability, complicating computational analysis. In a collaborative effort between our neuropathology and computational teams, we curated a novel high-quality dataset of 228 WSIs scanned at 40 \(\times \) magnification, from which 75,720 hematoxylin and eosin (H&E) patches were systematically extracted across six WHO-recognized glioma subtypes. Each patch was paired with an anonymized, neuropathologist authored diagnostic description, enabling clinically reliable image and text alignment. Building on this foundation, we present a novel vision–language contrastive framework that amalgamates a Vision Transformer (ViT-B/16) for morphological encoding with BioClinicalBERT for clinical text representation, jointly optimized in a shared embedding space via symmetric InfoNCE loss. The model achieved a Recall@10 of 55%, markedly outperforming the CLIP baseline, yet leaves scope for improvement in text-to-image and image-to-image retrieval tasks in neuropathology. Qualitative analyses further demonstrate the ability to capture fine-grained morphologic linguistic correspondences and establish a benchmark for multimodal retrieval in computational neuropathology.