<p>Recent advances in learned image compression (LIC) have demonstrated superior performance over traditional methods but often require training and storage of multiple models to handle different bitrate settings. In this paper, we propose the Uniform Spatial-Frequency Residual Bottleneck Modulation Adapter (U-SFRB), a plug-and-play, adapter-based framework for variable rate image compression that significantly reduces training and storage overhead. Our method freezes the backbone network and only trains lightweight adapters—Spatial-Frequency Residual Bottleneck Adapters (SFRBs)—to achieve rate adaptability. By inserting multiple SFRBs in parallel, our approach enables a single model to support a wide range of bitrates. Unlike prompt-based methods restricted to transformer architectures, our approach is compatible with both CNN- and transformer-based compression models. Experimental results on the Kodak and CLIC datasets show that our method achieves competitive rate-distortion performance compared to state-of-the-art variable rate compression approaches, with the advantage of lower training complexity and better model flexibility.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Variable rate compression with Uniform Spatial-Frequency Residual Bottleneck Adapter for learned image compression

  • Ran Wang,
  • Yongqiang Wang,
  • Heming Sun,
  • Jiro Katto

摘要

Recent advances in learned image compression (LIC) have demonstrated superior performance over traditional methods but often require training and storage of multiple models to handle different bitrate settings. In this paper, we propose the Uniform Spatial-Frequency Residual Bottleneck Modulation Adapter (U-SFRB), a plug-and-play, adapter-based framework for variable rate image compression that significantly reduces training and storage overhead. Our method freezes the backbone network and only trains lightweight adapters—Spatial-Frequency Residual Bottleneck Adapters (SFRBs)—to achieve rate adaptability. By inserting multiple SFRBs in parallel, our approach enables a single model to support a wide range of bitrates. Unlike prompt-based methods restricted to transformer architectures, our approach is compatible with both CNN- and transformer-based compression models. Experimental results on the Kodak and CLIC datasets show that our method achieves competitive rate-distortion performance compared to state-of-the-art variable rate compression approaches, with the advantage of lower training complexity and better model flexibility.