MCFM: An efficient multi-scale feature mixing network for image super-resolution
摘要
Modern CNN architectures can compete with Transformers in various vision tasks. Therefore, we propose a Multi-scale Chunked Feature Mixing Network (MCFM) for efficient image super-resolution. Our approach combines re-parameterization with Incremental Weight Optimization, maintaining multi-branch capability during training while ensuring single-branch efficiency at inference. The core of MCFM is the Multi-Scale Feature Decomposition (MSFD) module, which employs a three-tier strategy: MBRConvS captures fine-grained local features through multi-branch structures, MBRConvM extracts medium-scale contextual information, and decomposed strip convolutions efficiently model global spatial relationships. To enhance feature representation, we introduce the Gated Spatial Attention Unit (GSAU) that integrates spatial attention with gating mechanisms while reducing computational complexity. Additionally, the Hierarchical Dual-Path Attention (HDPA) mechanism realizes adaptive feature selection through collaborative spatial and global attention paths. Experimental results on benchmark datasets show that MCFM outperforms state-of-the-art methods by 0.2 dB on Urban100