Pricing strategies in mobile crowdsensing: an enhanced MAPPO approach using a behavior network
摘要
Mobile crowdsensing involves assigning multiple tasks around points of interest to mobile users (MUs) for execution. Developing an optimal task allocation strategy is crucial for the entire system, as it directly impacts the benefits of stakeholders. Leveraging recent advancements in multi-agent reinforcement learning (MARL), which have demonstrated unique advantages in simulating complex interactions among multiple agents, we propose a pricing strategy based on an improved behavior network and multi-agent proximal policy optimization (MAPPO) algorithm. Specifically, we formulate the problem as a multi-leader multi-follower Stackelberg game, and then apply MAPPO, a MARL technique which employs centralized training and decentralized execution, to solve this game. To better capture complex sequential input information and achieve superior behavior strategies, we integrate an attention mechanism with a gated recurrent unit (GRU) network into the actor network, forming a MARL algorithm with an improved behavior network, termed GRU-and-Attention-based MAPPO (GA-MAPPO). Simulation results demonstrate that the proposed GA-MAPPO algorithm is effective compared with baseline approaches. It can learn an optimal pricing strategy that maximizes the benefits of Task Initiators (TIs) and guides TIs in pricing MUs effectively.