Joint Source-Channel Coding with Large Language Model: A Vibrotactile Example
摘要
Recent advancements in tactile and communication technologies have created new opportunities for immersive virtual reality and remote-control applications. However, a key challenge is developing communication systems that are both efficient and adaptable to changing channel conditions. Achieving seamless, real-time tactile feedback requires advanced encoding strategies that adjust to dynamic factors such as bandwidth, latency, and packet loss. To address these challenges, this paper leverages the capabilities of the Large Language Model (LLM) ChatGPT 4.0, positioning it as an intelligent decision-making component in joint source-channel coding (JSCC) for vibrotactile data transmission. The model can process and respond to real-time channel variations, enabling adaptive encoding decisions. This optimization aligns with system conditions to enhance responsiveness and adaptability. This approach also simplifies development and maintenance while improving scalability and transmission efficiency. Furthermore, we propose a novel deep JSCC framework that integrates modules for semantic extraction, encoding, decoding, and channel feedback. The semantic extraction module effectively captures the essential features of vibrotactile data, while the channel adaptation module ensures resilience to noise. Comparative experiments on the IEEE P1918.1.1 haptic codec task force dataset demonstrate that model outperforms traditional separate communication schemes and existing JSCC methods. This study offers a new perspective on integrating LLM with JSCC, advancing tactile communication technology for next-generation immersive experiences.