Smart Grid Real-time Pricing for Multitype Users: A Multi-agent DQL-MHA-PER Algorithm for Welfare Equilibrium
摘要
This paper proposes a multi-agent double Q-learning algorithm integrated with multi-head attention and prioritized experience replay (DQL-MHA-PER) to address the welfare equilibrium problem in real-time pricing for multitype users in smart grids. By constructing independent decision models for power suppliers and three types of users, the framework achieves welfare equilibrium while accounting for electricity price fluctuations, renewable energy integration, and grid constraints. The introduced multi-head attention mechanism captures the non-linear correlations among dynamic electricity prices, user utility preferences, and load patterns, thereby significantly enhancing strategy precision. Prioritized experience replay accelerates the learning efficiency in sparse-reward scenarios, and double Q-learning mitigates the problem of value function overestimation. The model’s effectiveness is verified through simulation experiments under different pricing methods and volatility scenarios. In the algorithm comparison section, numerical comparisons across 360 experimental sets show that DQL-MHA-PER outperforms Q-learning, double Q-learning, and attention-based double Q-learning in both user welfare improvement and stability: its average welfare increases by 28%, 23.8%, and 17.6% respectively, while its standard deviation decreases by 18.8%, 16.1%, and 0.6% respectively. These findings not only validate the superiority of the proposed DQL-MHA-PER algorithm in dynamic pricing and welfare equilibrium scenarios but also provide practical guidance for the design of user-centric real-time pricing mechanisms and the optimization of multi-stakeholder collaborative strategies in smart grids.