<p>Vision transformers (ViTs) achieve state-of-the-art accuracy in recognition and segmentation, yet their self-attention makes decisions difficult to interpret. We call, and will refer to our method as ViTMix. Unlike prior CNN oriented ensemble explainers, ViTMix is tailored to ViTs and transfers to medical data with only brief target-domain fine tuning. Single-method explainers, Gradient Saliency, Grad-CAM, Layer-wise Relevance Propagation (LRP), or Attention Rollout, capture only partial evidence and often yield noisy heat maps. We introduce a model-agnostic post hoc fusion framework that supports combining multiple attribution maps via element wise multiplication and geometric mean. The mixed map keeps pixels highlighted by multiple methods while suppressing isolated artifacts, a behavior we motivate analytically via the Pigeonhole Principle. ViTMix operates on any ViT classifier; its sole requirement is a model that outputs class logits and-aside from the light fine tuning used on PH<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(^{2}\)</EquationSource> </InlineEquation>-generalizes from natural images to dermoscopic scans. On ImageNet, the LRP&#xa0;+&#xa0;Rollout fusion increases IoU by 10.66 points (38.53–49.19) and F1 by 10.19 points (53.18–63.37) while slightly reducing deletion-AUC by 0.01 (0.44–0.43); on Pascal VOC it increases IoU by 4.30 points (36.31–40.61) and F1 by 4.54 points (52.04–56.58), with deletion-AUC 0.55. Applied to the PH<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(^{2}\)</EquationSource> </InlineEquation> medical set, the same fusion attains 64.5 % IoU and 76.7 % F1, surpassing a vanilla ViT by more than 20 % relative. Student annotations further prove our point, with Cohen’s <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\kappa =0.79\)</EquationSource> </InlineEquation> and a Dice overlap of 0.85 indicating substantial agreement with human perception. These results show that mixing complementary XAI signals yields clearer, more faithful explanations and reliable weak segmentation masks without architectural changes or additional supervision. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From explanation to unsupervised segmentation: fusion of multiple explanation maps for vision transformers

  • Eduard Hogea,
  • Darian M. Onchis,
  • Ana Coporan,
  • Adina Magda Florea,
  • Codruta Istin

摘要

Vision transformers (ViTs) achieve state-of-the-art accuracy in recognition and segmentation, yet their self-attention makes decisions difficult to interpret. We call, and will refer to our method as ViTMix. Unlike prior CNN oriented ensemble explainers, ViTMix is tailored to ViTs and transfers to medical data with only brief target-domain fine tuning. Single-method explainers, Gradient Saliency, Grad-CAM, Layer-wise Relevance Propagation (LRP), or Attention Rollout, capture only partial evidence and often yield noisy heat maps. We introduce a model-agnostic post hoc fusion framework that supports combining multiple attribution maps via element wise multiplication and geometric mean. The mixed map keeps pixels highlighted by multiple methods while suppressing isolated artifacts, a behavior we motivate analytically via the Pigeonhole Principle. ViTMix operates on any ViT classifier; its sole requirement is a model that outputs class logits and-aside from the light fine tuning used on PH \(^{2}\) -generalizes from natural images to dermoscopic scans. On ImageNet, the LRP + Rollout fusion increases IoU by 10.66 points (38.53–49.19) and F1 by 10.19 points (53.18–63.37) while slightly reducing deletion-AUC by 0.01 (0.44–0.43); on Pascal VOC it increases IoU by 4.30 points (36.31–40.61) and F1 by 4.54 points (52.04–56.58), with deletion-AUC 0.55. Applied to the PH \(^{2}\) medical set, the same fusion attains 64.5 % IoU and 76.7 % F1, surpassing a vanilla ViT by more than 20 % relative. Student annotations further prove our point, with Cohen’s \(\kappa =0.79\) and a Dice overlap of 0.85 indicating substantial agreement with human perception. These results show that mixing complementary XAI signals yields clearer, more faithful explanations and reliable weak segmentation masks without architectural changes or additional supervision.