SCST: Spatial Consistent Swin Transformer for Multi-focus Biomedical Microscopic Image Fusion
摘要
This paper studies multi-focus biomedical microscopic image fusion. Recently, deep convolutional neural network (CNN) based approaches have achieved promising performance in multi-focus image fusion, but most of them cannot obtain spatially continuous results, especially in smooth regions and edges between focused and defocused regions. To address this issue, we propose a novel method termed Spatial Consistent Swin Transformer, which utilizes Residual Swin Transformer combined with Squeeze-and-Excitation block to make the fusion results to be spatially consistent. To facilitate comprehensive model training, we create a substantial dataset of multi-focus biomedical microscopic images, primarily derived from ARCH dataset. To avoid the defocus spread effect around a focus/defocus boundary, we incorporate a trainable self-guided filtering module, which can further enhance spatial consistency in the predicted focus map and also eliminate the need for post-processing. Experimental results on both the popular EDoF Fraunhofer dataset and our self-collected datasets show that the proposed method outperform the state-of-the-art approaches in terms of both standard quantitative image fusion metrics and visual quality.