<p>Recently, Transformer-based techniques have demonstrated impressive effectiveness across various high- and low-level vision tasks by leveraging the self-attention mechanism for feature extraction. However, using self-attention is computationally expensive for applications with low computational resources. To solve the Transformer problem, we propose the convolution network (ConvNextN) based on the original convolution network (ConvNextv2) that is used for high-level vision. The ConvNextN network has the Transformer and convolution neural networks (CNNs) merits with only convolution layers. The ConvNeXtv2 contains the depthwise convolution, then pointwise convolution, GRN, and one pointwise convolution. The ConvNextN is based on using the ConvNextv2 group as the backbone with 3 <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2024_20492_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 3 convolution and enhanced spatial attention. This ConvNextG is based on the ConvNextv2 block (ConvNextB) as a main block in a Transformer-style block. This ConvNextB is built using the ConvNextv2 with layer norm and depthwise multi-layer perceptron. The global response normalization is used inside the ConvNextv2 to enhance inter-channel feature competition. This model explores the impact of using the Transformer style on the traditional convolution layers. Finally, this model attained good performance in multiple super-resolution benchmarks. Our model achieved 0.03 dB better PSNR than LWSwinIR in the Urban100 dataset with an efficient number of parameters, Mult-adds, and runtime. Also, it is clear that our model has 0.05 dB enhancement in PSNR compared to LWSwinIR for the Set5 dataset at the scale of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2024_20492_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 2.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-style convolution network for lightweight image super-resolution

  • Garas Gendy,
  • Nabil Sabor

摘要

Recently, Transformer-based techniques have demonstrated impressive effectiveness across various high- and low-level vision tasks by leveraging the self-attention mechanism for feature extraction. However, using self-attention is computationally expensive for applications with low computational resources. To solve the Transformer problem, we propose the convolution network (ConvNextN) based on the original convolution network (ConvNextv2) that is used for high-level vision. The ConvNextN network has the Transformer and convolution neural networks (CNNs) merits with only convolution layers. The ConvNeXtv2 contains the depthwise convolution, then pointwise convolution, GRN, and one pointwise convolution. The ConvNextN is based on using the ConvNextv2 group as the backbone with 3 \(\times \) × 3 convolution and enhanced spatial attention. This ConvNextG is based on the ConvNextv2 block (ConvNextB) as a main block in a Transformer-style block. This ConvNextB is built using the ConvNextv2 with layer norm and depthwise multi-layer perceptron. The global response normalization is used inside the ConvNextv2 to enhance inter-channel feature competition. This model explores the impact of using the Transformer style on the traditional convolution layers. Finally, this model attained good performance in multiple super-resolution benchmarks. Our model achieved 0.03 dB better PSNR than LWSwinIR in the Urban100 dataset with an efficient number of parameters, Mult-adds, and runtime. Also, it is clear that our model has 0.05 dB enhancement in PSNR compared to LWSwinIR for the Set5 dataset at the scale of \(\times \) × 2.