Semantic Communication of Images Using Image Generation and Image Captioning Models
摘要
With the emergence of 6G networks and the new networking requirements, the existing networks are approaching the Shannon capacity limit. Hence, a new paradigm called ‘semantic communication’ is proposed in the literature. Semantic communication is the transmission of the meaning of a message between two intelligent agents. In this work, the problem of image semantic communication is considered. Most of the existing works in literature focus on the exact reconstruction of the transmitted image. Yet, to achieve full semantic communication, it is possible to focus on transmitting a similar image with the same meaning as the original image. This can be achieved by transmitting a text description of the image. An image captioning model, namely, the one-for-all (OFA) model is used to transform an image into its caption. The caption is transmitted through the network and at the receiver an image-to-text diffusion model, stability diffusion XL, is used to reconstruct a semantically similar image. The similarity between the transmitted and the received images is assessed using the textual similarity between the corresponding captions. The captions of the original and the generated images had a similarity above 0.7 using five text similarity measures, namely, ANLS, METEOR, ROUGE, BLEU and BERT similarity. This approach is shown to achieve a 99% reduction in the message size. The results emphasize the effect of using AI models to extract a general representation of a message and in turn consume less communication resources.