<p>Semantic segmentation faces significant challenges in domain generalization due to diverse environmental variations and complex scene structures. In this work, we propose ES-TQDM, a structure-aware and lightweight framework that enhances multi-scale adaptability and spatial perception by integrating two complementary attention modules: Efficient Multi-Scale Attention (EMA) and Simple Parameter-free Attention Module (SimAM). EMA captures global semantic coherence through hierarchical multi-scale fusion, while SimAM selectively refines fine-grained local details, particularly small objects and boundaries. Their complementary design enables ES-TQDM to consistently improve segmentation robustness across diverse domains and backbones (ViT-B and EVA02-L). Extensive experiments demonstrate up to a 2.13-point&#xa0;mIoU gain in challenging settings such as Cityscapes → BDD100K, with minimal computational overhead (only 0.05–0.06% GFLOPs increase). Ablation studies further confirm the synergy between EMA and SimAM in boosting cross-domain generalization. Our results highlight that targeted, structure-aware fusion of complementary attention mechanisms is an effective strategy for addressing semantic segmentation under domain shifts. The project page is available at <a href="https://github.com/V1perQAQ0821/ES-TQDM">https://github.com/V1perQAQ0821/ES-TQDM</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient semantic segmentation across domains: enhancing generalization with multi-scale and simple attention modules

  • Yuan Luo,
  • Junlei Chen,
  • Zerui Yao

摘要

Semantic segmentation faces significant challenges in domain generalization due to diverse environmental variations and complex scene structures. In this work, we propose ES-TQDM, a structure-aware and lightweight framework that enhances multi-scale adaptability and spatial perception by integrating two complementary attention modules: Efficient Multi-Scale Attention (EMA) and Simple Parameter-free Attention Module (SimAM). EMA captures global semantic coherence through hierarchical multi-scale fusion, while SimAM selectively refines fine-grained local details, particularly small objects and boundaries. Their complementary design enables ES-TQDM to consistently improve segmentation robustness across diverse domains and backbones (ViT-B and EVA02-L). Extensive experiments demonstrate up to a 2.13-point mIoU gain in challenging settings such as Cityscapes → BDD100K, with minimal computational overhead (only 0.05–0.06% GFLOPs increase). Ablation studies further confirm the synergy between EMA and SimAM in boosting cross-domain generalization. Our results highlight that targeted, structure-aware fusion of complementary attention mechanisms is an effective strategy for addressing semantic segmentation under domain shifts. The project page is available at https://github.com/V1perQAQ0821/ES-TQDM.