Multimodal-Guided Perceptual Image Compression via Joint Text and Audio
摘要
Balancing the rate-distortion-perception trade-off in image compression is a challenging task, as minimizing distortion often degrades perceptual quality of the reconstructed images. Existing perceptual image compression methods rely solely on text guidance, limiting their ability to leverage richer multimodal contexts. Since audio often coexists with text and images, utilizing such multimodal information provides a more comprehensive understanding for guiding image compression and improving perceptual quality. In this paper, we propose a Multimodal-Guided Image Compression network (MGIC) that leverages semantic information from both text and audio to enhance the perceptual quality of the reconstructed image. To effectively fuse multimodal features at both local and global scales, we design a Multimodal Feature Fusion Block (MFFB). This module enables multimodal features to optimize local details of the image while leveraging global semantics to perceive the overall content of the image. Extensive experiments show that at low bitrates, our method achieves better perceptual quality reconstruction with significantly lower bitrates, e.g., 0.6 × bits of HiFiC and 0.4 × bits of Bpg.