FPSNet: Focus-Perceptual-Semantic Full Flow Visual Redundancy Predicting for Camera Image
摘要
With the rapid popularization of electronic devices, the large amount of data generated by camera imaging poses a huge challenge to the limited storage capacity and communication bandwidth. Achieving higher compression ratios without sacrificing visual quality remains a fundamental challenge for image compression. In this paper, we propose a novel full flow bidirectional visual threshold estimation method for camera imaging perceptual compression. Specifically, we study the features from camera imaging to visual perception to semantic understanding, and characterize them with focus identification, perceptual distribution, and semantic segmentation respectively. We also carefully design feature extraction networks suitable for each feature type. In addition, we draw inspiration from the bidirectional perceptual mechanism of the human visual system and propose a feature extraction framework that adopts top-down and bottom-up methods. We further enhance our model by regulating and fusing bidirectional perceptual features through a gated decoding structure. Extensive experimental validation on benchmark datasets confirms that our FPSNet significantly improves the accuracy of visual redundancy prediction.