<p>Semantic segmentation involves assigning a class label to every pixel in an image and serves as a key technology in applications such as medical imaging, autonomous driving, and satellite image analysis. While existing deep learning models and transformer-based architectures have demonstrated outstanding performance, their high computational costs and memory demands limit their applicability in resource-constrained environments. To address these challenges, knowledge distillation (KD), which transfers knowledge from a high-performing teacher model to a lightweight student model, has emerged as an effective approach. However, existing logit-based KD methods have shown limitations in fully leveraging spatial context and structural information. Additionally, the distillation process often introduces computational and memory overhead, restricting its practicality. This study proposes a novel framework, Multi Spatial Projectors for Knowledge Distillation (MSPKD), to overcome these limitations. At its core, MSPKD utilizes lightweight <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="138_2025_1721_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(1 \times 1\)</EquationSource> </InlineEquation> convolutional layers as multi spatial projectors to maintain spatial integrity while effectively aligning the feature spaces between teacher and student models. The total loss function combines cross-entropy loss, standard KD loss, and projector-based KD loss to effectively support the student model’s learning. Experiments conducted on representative segmentation datasets—PASCAL VOC, Cityscapes, and CamVid—demonstrate that the proposed MSPKD framework achieves performance improvements in both mean intersection over union (mIoU) and pixel accuracy compared to Vanilla KD, while simultaneously proving its efficiency and practicality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MSPKD: multi spatial projectors for knowledge distillation in semantic segmentation

  • Yeongje Park,
  • Jeongwon Hwang,
  • Seung-Min Jeong,
  • Eui Chul Lee

摘要

Semantic segmentation involves assigning a class label to every pixel in an image and serves as a key technology in applications such as medical imaging, autonomous driving, and satellite image analysis. While existing deep learning models and transformer-based architectures have demonstrated outstanding performance, their high computational costs and memory demands limit their applicability in resource-constrained environments. To address these challenges, knowledge distillation (KD), which transfers knowledge from a high-performing teacher model to a lightweight student model, has emerged as an effective approach. However, existing logit-based KD methods have shown limitations in fully leveraging spatial context and structural information. Additionally, the distillation process often introduces computational and memory overhead, restricting its practicality. This study proposes a novel framework, Multi Spatial Projectors for Knowledge Distillation (MSPKD), to overcome these limitations. At its core, MSPKD utilizes lightweight \(1 \times 1\) convolutional layers as multi spatial projectors to maintain spatial integrity while effectively aligning the feature spaces between teacher and student models. The total loss function combines cross-entropy loss, standard KD loss, and projector-based KD loss to effectively support the student model’s learning. Experiments conducted on representative segmentation datasets—PASCAL VOC, Cityscapes, and CamVid—demonstrate that the proposed MSPKD framework achieves performance improvements in both mean intersection over union (mIoU) and pixel accuracy compared to Vanilla KD, while simultaneously proving its efficiency and practicality.