To survive in a changing world, animals often need to suppress an obsolete behavior and acquire a new one. This process is known as reversal learning. The neural mechanisms underlying RL in spatial navigation have received limited attention and it remains unclear what neural mechanisms maintain behavioral flexibility. Here, we extend a closed-loop simulator of spatial navigation and learning, based on spiking neural networks [7]. In this model, activity of place cells and boundary cells are fed as inputs to action selection neurons, which drive the movement of the agent. Upon reaching the goal, behavior is reinforced with spike-timing-dependent plasticity (STDP) coupled with an eligibility trace which marks synaptic connections for future reward-based updates. We model a task with an ABA design, where the goal is switched between two locations A and B after 10 trials. Agents using symmetric STDP excel initially on finding target A, but fail to find target B after the goal switch, persevering on target A. Either using asymmetric STDP, using many small place fields, or injecting short noise pulses to action selection neurons are effective in driving spatial exploration, which ultimately leads to finding target B. However, this flexibility comes at the price of slower learning and lower performance. Our work shows three examples of neural mechanisms that achieve flexibility at the behavioral level, each with different characteristic costs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Cost of Behavioral Flexibility: Reversal Learning Driven by a Spiking Neural Network

  • Behnam Ghazinouri,
  • Sen Cheng

摘要

To survive in a changing world, animals often need to suppress an obsolete behavior and acquire a new one. This process is known as reversal learning. The neural mechanisms underlying RL in spatial navigation have received limited attention and it remains unclear what neural mechanisms maintain behavioral flexibility. Here, we extend a closed-loop simulator of spatial navigation and learning, based on spiking neural networks [7]. In this model, activity of place cells and boundary cells are fed as inputs to action selection neurons, which drive the movement of the agent. Upon reaching the goal, behavior is reinforced with spike-timing-dependent plasticity (STDP) coupled with an eligibility trace which marks synaptic connections for future reward-based updates. We model a task with an ABA design, where the goal is switched between two locations A and B after 10 trials. Agents using symmetric STDP excel initially on finding target A, but fail to find target B after the goal switch, persevering on target A. Either using asymmetric STDP, using many small place fields, or injecting short noise pulses to action selection neurons are effective in driving spatial exploration, which ultimately leads to finding target B. However, this flexibility comes at the price of slower learning and lower performance. Our work shows three examples of neural mechanisms that achieve flexibility at the behavioral level, each with different characteristic costs.