Long Short-Term Deterministic Policy Gradient for Joint Optimization of Computational Offloading and Resource Allocation in MEC
摘要
Mobile Edge Computing (MEC) is regarded as a promising paradigm for reducing service latency in Mobile users data processing by providing computing resources at the network edge. Existing deep reinforcement learning (DRL) algorithms struggle to effectively handle the joint optimization of computational offloading and resource allocation (JCORA). To overcome this challenge, we propose a Long Short-Term Deterministic Policy Gradient (LSTDPG) approach to tackle JCORA. Building upon the Deep Deterministic Policy Gradients (DDPG) algorithm, LSTDPG incorporates two key features. Firstly, it utilizes a Temporal Attention Network composed of Long Short-Term Memory (LSTM) networks, which facilitates high-quality state representation and function approximation. Secondly, an Episode-Based Prioritized Experience Replay (ePER) method is introduced to expedite and stabilize the convergence of model training. Experimental results demonstrate that the proposed LSTDPG outperforms several state-of-the-art DRL agents in terms of task completion time and energy consumption.