GATR: a graph attention deep reinforcement learning approach with variance-sensitive rewards for rpl routing optimization
摘要
The Internet of Things (IoT) has become the core infrastructure supporting intelligent perception and interconnection. However, the dynamic network topology, limited resources, and partial observability pose significant challenges to traditional IoT routing algorithms in balancing energy efficiency, reliability, and stability optimally. Therefore, this paper proposes a graph attention deep reinforcement learning-based routing optimization algorithm (GATR), which can optimize the data transmission performance through multi-agent collaborative decision-making. GATR integrates the Actor-Critic architecture with centralized training and decentralized execution (CTDE), and innovatively introduces a topology-aware dual-channel graph modeling mechanism to dynamically aggregate neighborhood information while explicitly encoding physical interference relationships, thereby mitigating partial observability constraints. For resource-constrained terminal devices, GATR designs a lightweight three-layer convolutional Actor network architecture, and introduces contrastive learning to enhance the feature representation of the routing decision model. Furthermore, a variance-sensitive multi-objective reward function formulation is proposed to dynamically balance energy consumption, reliability, and stability through adaptive weight adjustment. Simulation results demonstrate that GATR achieves significant improvements in energy efficiency, network lifetime, and convergence performance compared with RARL, DDQN, DADR, CARL, and MRHOF algorithms, exhibiting excellent environmental adaptability and engineering application potential.