<p>Data-driven image segmentation methods frequently encounter challenges related to object size imbalance, particularly in cases where small targets occupy only a tiny proportion of the image. Furthermore, the theoretical foundation underlying how these methods learn object size remains poorly understood. To address this problem, we study data-driven image segmentation from the perspective of semi-dual optimal transport and introduce image-specific volume prior into the activation layer of the segmentation neural network. In this framework, the neural network logits are interpreted to the transport cost, and the bias term in the activation operator is associated with the dual variable that controls the distribution of target size. Based on this theoretical discovery, we develop a nonlinear activation mechanism, termed <i>VP-Sparsemax</i>, in image segmentation neural networks with volume priors from smooth approximations of the max operator in the semi-dual optimal transport problem. Different from existing loss-based modifications, the proposed nonlinear activation module incorporates both volume prior and spatial information, and can be embedded into the segmentation networks for both training and inference. Theoretical analysis and numerical results show that the proposed method can improve the ability of the networks to learn object size. Experiments on several real-world datasets and some representative segmentation network backbones such as U-Net, DeepLabV3+, and SAM, show that the proposed method consistently outperforms the existing related methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Segmenting Objects with Imbalanced Sizes via Smooth and Sparse Dual Optimal Transport

  • Mengqi Ding,
  • Gangxuan Zhou,
  • Xue-Cheng Tai,
  • Li Cui,
  • Jun Liu

摘要

Data-driven image segmentation methods frequently encounter challenges related to object size imbalance, particularly in cases where small targets occupy only a tiny proportion of the image. Furthermore, the theoretical foundation underlying how these methods learn object size remains poorly understood. To address this problem, we study data-driven image segmentation from the perspective of semi-dual optimal transport and introduce image-specific volume prior into the activation layer of the segmentation neural network. In this framework, the neural network logits are interpreted to the transport cost, and the bias term in the activation operator is associated with the dual variable that controls the distribution of target size. Based on this theoretical discovery, we develop a nonlinear activation mechanism, termed VP-Sparsemax, in image segmentation neural networks with volume priors from smooth approximations of the max operator in the semi-dual optimal transport problem. Different from existing loss-based modifications, the proposed nonlinear activation module incorporates both volume prior and spatial information, and can be embedded into the segmentation networks for both training and inference. Theoretical analysis and numerical results show that the proposed method can improve the ability of the networks to learn object size. Experiments on several real-world datasets and some representative segmentation network backbones such as U-Net, DeepLabV3+, and SAM, show that the proposed method consistently outperforms the existing related methods.