Robust Policy Gradient Q-Learning with Particle Swarm Optimized Resource Allocation for 5G D2D Communication Networks
摘要
For effective resource allocation in device-to-device (D2D) networks in 5G noisy environments, we propose a robust policy gradient Q-learning algorithm with particle swarm optimization (PSO) in this paper. The suggested algorithm, dubbed RPSO-QPG, combines Q-learning, policy gradient, and PSO in order to enhance the efficiency of resource allocation in D2D networks. The algorithm uses a reliable policy gradient optimization technique to adapt to the chaotic and dynamic environment of 5G networks. Using simulations, we assess the proposed algorithm’s performance in terms of convergence rate, memory needs, and system throughput. Our findings demonstrate that the RPSO-QPG algorithm performs better in terms of system throughput and convergence speed, while also requiring less memory, than existing Q-learning and Q-learning with policy gradient algorithms. We come to the conclusion that the RPSO-QPG algorithm is a promising strategy for effective resource allocation in 5G noisy environments for D2D networks.