Deep Image Coding in the Fractional Wavelet Transform Domain based on High-Frequency Sub-bands Prediction
摘要
This paper presents an image coding method using deep convolutional neural networks and the fractional wavelet transform (FrWT) algorithm. FrWT requires less memory than traditional discrete wavelet transform (DWT) methods. During the image transformation from the pixel domain to the wavelet domain, one low-frequency sub-band (LF-sub-band) and three high-frequency sub-bands (HF-sub-bands) are generated. The LF-sub-band is used to predict each HF-sub-band, reducing redundancy between them. The sub-bands are then fed into separate auto-encoders for encoding. Additionally, a conditional probability model is used to estimate the context-dependent prior probability of the encoded codes, improving entropy coding efficiency. A joint framework is used to train the auto-encoders and conditional probability model. Experimental results show that the proposed approach outperforms JPEG, JPEG2000, BPG, and some other neural network-based image coding techniques in terms of multi-scale structural similarity index measure (MS-SSIM). The proposed method generates better visual quality with more precise details and textures due to increased high-frequency prediction. The HF images exhibit average gains of 0.0011, 0.0027, and 0.0049 in MS-SSIM under low, middle, and high bitrates, respectively.