Weakly supervised segmentation has the potential to greatly reduce the annotation effort for training segmentation models for small structures such as hyper-reflective foci (HRF) in optical coherence tomography (OCT). However, most weakly supervised methods either involve a strong downsampling of input images, or only achieve localization at a coarse resolution, both of which are unsatisfactory for small structures. We propose a novel framework that increases the spatial resolution of a traditional attention-based multiple instance learning (MIL) approach by using layer-wise relevance propagation (LRP) to prompt the segment anything model (SAM 2), and increases recall with iterative inference. Moreover, we demonstrate that replacing MIL with a compact convolutional transformer (CCT), which adds a positional encoding, and permits an exchange of information between different regions of the OCT image, leads to a further and substantial increase in segmentation accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weakly Supervised Segmentation of Hyper-reflective Foci with Compact Convolutional Transformers and SAM 2

  • Olivier Morelle,
  • Justus Bisten,
  • Maximilian WM. Wintergerst,
  • Robert P. Finger,
  • Thomas Schultz

摘要

Weakly supervised segmentation has the potential to greatly reduce the annotation effort for training segmentation models for small structures such as hyper-reflective foci (HRF) in optical coherence tomography (OCT). However, most weakly supervised methods either involve a strong downsampling of input images, or only achieve localization at a coarse resolution, both of which are unsatisfactory for small structures. We propose a novel framework that increases the spatial resolution of a traditional attention-based multiple instance learning (MIL) approach by using layer-wise relevance propagation (LRP) to prompt the segment anything model (SAM 2), and increases recall with iterative inference. Moreover, we demonstrate that replacing MIL with a compact convolutional transformer (CCT), which adds a positional encoding, and permits an exchange of information between different regions of the OCT image, leads to a further and substantial increase in segmentation accuracy.