An adaptive online reinforcement learning approach for optimizing computation offloading in connected vehicle networks
摘要
Processing computational tasks in intelligent applications installed on vehicles poses a significant challenge in dynamic intelligent transportation networks. The limited computational and storage resources of vehicles make handling heavy computational tasks difficult. Leveraging V2V communication technology and adaptive intelligent solutions can facilitate faster processing, enhance computational efficiency, and reduce failures in delivering task results. This paper proposes a model-free, adaptive online reinforcement learning approach based on the Q-learning algorithm for computation offloading in connected vehicle networks (CVNs). This approach leverages feedback derived from Q-value estimates at each moment and long-term rewards to refine its decision-making process for selecting the most suitable vehicle for task offloading. Consequently, it improves its baseline policy to ultimately achieve long-term optimization objectives. A key advantage of this online approach is its reduced reliance on intensive computations and large datasets compared to traditional learning methods, along with its rapid adaptability to environmental changes. Experimental results demonstrate that the proposed approach outperforms priority, balanced, and MMD algorithms with improvements of 6.09%, 2.24%, and 0.80%, respectively, in computational efficiency. Additionally, the method achieves a substantial increase of 2.23 × 107 and 1.14 × 106 in total long-term rewards and good action selection, respectively, highlighting its significant effectiveness. The proposed algorithm, utilizing feedback derived from Q-values, reduces the switching rate for policy improvement by 3.71%, demonstrating the algorithm's enhancement and policy optimization as the arrival rate increases.