Policy-Based Adaptive Interference Decision-Making in Complex Dynamic Environments
摘要
Multifunctional radar (MFR) is highly adaptable and places more emphasis on the cognitive abilities of the system. The continuous improvement of the flexibility and performance potential of MFR systems has gradually widened the technical gap with electronic jamming equipment. This paper focuses on the problem of how a jammer can autonomously and reliably learn optimal jamming strategy in complex environments by interacting with MFR that possesses strong adaptive capabilities and flexible mode conversion functions. We propose applying the policy-based reinforcement learning (RL) algorithm to adaptive jamming decision-making (JDM). This algorithm no longer indirectly explores the optimal strategy by maintaining the value function in RL. Instead, it parameterizes the strategy and directly learns it, iteratively updating the strategy through the interaction between the jammer and MFR. Finally, two typical deep Q network (DQN) and double DQN (DDQN) algorithms were compared and analyzed through simulation experiments. The simulation results demonstrate that this algorithm not only achieves higher returns than the value-based algorithms but also exhibits faster convergence speed and stable performance. These findings suggest that the algorithm possesses better adaptability in complex dynamic environments.