Actor-Critic Spiking Neural Network with RSTDP Actor Learning and TD-LTP Critic Learning
摘要
An actor-critic spiking neural network is presented, where the critic is trained on base of temporal difference long-term potentiation, and the actor network is trained on base of reward-modulated spike timing-dependent plasticity. The proposed network achieves competitive performance on the Acrobot benchmark. This result is a preliminary step towards reinforcement learning of spiking networks that could be implemented in energy-efficient neuromorphic hardware.