Abstract <p>Pulse neural networks, suitable for hardware implementation based on memristors, are promising for robotics due to their energy efficiency. However, reinforcement learning algorithms using such networks remain poorly understood. One of the key motivations for using memristors as network weights is, in addition to energy efficiency, their ability to learn (change conductivity) in real time by superimposing voltage pulses from pre- and postsynaptic signals. This article presents the results of numerical simulation of a spiking neural network (SNN) with memristive synaptic connections, approximately solving an optimal control problem using trace variables to change weights, allowing one to approach reinforcement learning in real time. The fundamental possibility of such training in the task of holding a pole on a moving platform is shown, a comparison of various reward functions is given, and assumptions are made about ways to increase the effectiveness of this approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reinforcement Learning of Spiking Neural Networks Using Trace Variables for Synaptic Weights with Memristive Plasticity

  • V. A. Kulagin,
  • A. N. Matsukatova,
  • V. V. Ryl’kov,
  • V. A. Demin

摘要

Abstract

Pulse neural networks, suitable for hardware implementation based on memristors, are promising for robotics due to their energy efficiency. However, reinforcement learning algorithms using such networks remain poorly understood. One of the key motivations for using memristors as network weights is, in addition to energy efficiency, their ability to learn (change conductivity) in real time by superimposing voltage pulses from pre- and postsynaptic signals. This article presents the results of numerical simulation of a spiking neural network (SNN) with memristive synaptic connections, approximately solving an optimal control problem using trace variables to change weights, allowing one to approach reinforcement learning in real time. The fundamental possibility of such training in the task of holding a pole on a moving platform is shown, a comparison of various reward functions is given, and assumptions are made about ways to increase the effectiveness of this approach.