PPO-Based Adaptive Waveform Design for Intermittent Sampling and Replaying Jamming
摘要
To address the challenge of jamming against dynamically changing radar waveforms, this paper proposes an adaptive intermittent sampling and replaying jamming waveform design method based on the Proximal Policy Optimization (PPO) algorithm. By constructing a reward function centered on the jamming-to-signal ratio (JSR), the method guides the agent to adaptively determine the sampling and replaying parameters of the jamming signal, thereby achieving dynamic optimization of the jamming strategy. A multi-channel time-frequency feature representation is designed as the state input, a discrete parameterized action space is constructed, and ResNet18 is employed for feature extraction. Simulation results demonstrate that, compared with DQN and A2C, the proposed method achieves significantly better jamming performance and faster convergence, verifying its effectiveness and adaptability in complex radar environments.