Semantic Communications with Zero-Shot Learning for Multi-modal Transmission
摘要
Semantic communications (SCs) have become pivotal in ensuring efficient and accurate data transmission, especially in complex scenarios involving object recognition. This paper explores the integration of multi-modal data streams and zero-shot learning (ZSL) to enhance user interaction and system adaptability in virtual environments. We introduce a generative algorithm equipped with a multi-modal alignment loss function, improving the adaptability and efficiency of SC systems by enabling them to process and integrate real-time data with high precision. Furthermore, we propose semantic metrics for evaluating the performance of these systems, particularly their ability to handle previously unknown data types. Finally, practical experiments are conducted to validate the effectiveness of the proposed scheme.