Vision and LiDAR multi-modal fusion beam prediction method for millimeter-wave communication system
摘要
Future vehicle-connected communication systems increasingly rely on millimeter-wave (mmWave) and terahertz (THz) communication technologies to establish reliable connections. While high-frequency band communications offer the advantages of high data rates and low latency, the beams are susceptible to interference from obstacles, which leads to signal attenuation and multipath effects, affecting the reliability of traditional unimodal beam prediction. Therefore, the use of multi-modal sensor data such as Vision and LiDAR is crucial. However, the adoption of these techniques often results in a higher beam training overhead. To solve this problem, we propose a new multi-modal deep learning framework that utilizes a Multi-modal Group Attention Transformer (MGA-Transformer) for sensor-assisted beam prediction. Specifically, our method first preprocesses the collected visual and LiDAR data, utilizes a Grouped-query Attention mechanism for multi-scale feature fusion, and utilizes time mixing to enhance the temporal correlation of multi-modal data. Finally, the fused multi-modal features are fed into the Multilayer Perceptron (MLP) network to achieve fast and accurate beam prediction. In four different communication scenarios, our proposed scheme shows an overall prediction accuracy of 91.4%, a 10% improvement in the Top-1 accuracy compared with single-mode beam prediction. Because the sensor data are significantly affected by the environment, the performance of the multi-modal model is degraded when there is only a single-mode input. We adopted the Moe architecture to provide an independent processing unit for each mode and verified that the multi-modal model can effectively handle situations when there is only image data input. The solution effectively reduces the beam training overhead while achieving a reasoning speed of 43.47 FPS, providing highly reliable and low-latency communication support in highly mobile environments.