Accurate rectal cancer grading requires complementary CT images and pathological text. However, applying vision-language models faces challenges: limited disease category diversity and image-text mismatches hinder large-scale, one-to-one pair construction for contrastive learning. Additionally, complex CT anatomy with irrelevant tissues interferes with tumor identification, a problem general-purpose models struggle with, necessitating a specialized rectal-focused approach. This paper proposes a rectal tumor grading method based on weak alignment of image and pathological text features. We leverage pathological text to guide image features toward tumor category prototypes, avoiding strict one-to-one pairing. Large language models refine medical text, while segmentation models generate foreground channels to isolate the rectal region, improving feature extraction. A cross-modal attention mechanism further enhances alignment, leading to more accurate tumor grading, with experimental results showing a 2.3% accuracy improvement over existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Integration Based on Weak Alignment for Rectal Tumor Grading

  • Hongwu Liu,
  • Shouhong Wan,
  • Chenyang Qiu,
  • Bingbing Zou,
  • Wanqin Wang,
  • Risheng Xie,
  • Peiquan Jin

摘要

Accurate rectal cancer grading requires complementary CT images and pathological text. However, applying vision-language models faces challenges: limited disease category diversity and image-text mismatches hinder large-scale, one-to-one pair construction for contrastive learning. Additionally, complex CT anatomy with irrelevant tissues interferes with tumor identification, a problem general-purpose models struggle with, necessitating a specialized rectal-focused approach. This paper proposes a rectal tumor grading method based on weak alignment of image and pathological text features. We leverage pathological text to guide image features toward tumor category prototypes, avoiding strict one-to-one pairing. Large language models refine medical text, while segmentation models generate foreground channels to isolate the rectal region, improving feature extraction. A cross-modal attention mechanism further enhances alignment, leading to more accurate tumor grading, with experimental results showing a 2.3% accuracy improvement over existing methods.