错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pareto Deterministic Policy Gradients and Its Application in 6G Networks

  • Zhou Zhou,
  • Yan Xin,
  • Hao Chen,
  • Charlie Zhang,
  • Lingjia Liu,
  • Kai Yang

摘要

In this chapter, we introduce a reinforcement learning (RL)-based approach to jointly optimize cell load balance and network throughput as a potential AI/ML-based use case for sixth-generation cellular systems (6G), where inter-cell handover and massive MIMO antenna tilting are configured as the RL policy to learn. Our rationale behind using RL is to circumvent the challenges of analytically modeling user mobility and network dynamics. We integrate vector rewards into multiple value networks and conduct RL action via a separate policy network. We name this method as Pareto deterministic policy gradients (PDPG). It is an actor-critic, model-free, and deterministic policy algorithm which can handle the coupling objectives with the following two merits: (1) It solves the optimization via leveraging the degree of freedom of vector reward as opposed to choosing handcrafted scalar reward; (2) cross-validation over multiple policies can be significantly reduced. To be self-contained, an ideal static optimization-based brute-force search solver is included as the benchmark method. The comparison shows that the RL approach performs as well as this ideal strategy, though the former one is constrained with limited environment observations and lower action frequency, whereas the latter one has full access to the user mobility.