In most memristive neural network circuits based on operant conditioning, the agent’s tendency towards certain behaviors is simply reflected through changes in synaptic weight. No specific analysis has been conducted on the changes in the tendency of intelligent agent behavior. Therefore, an operant conditioning circuit based on memristors and Q-learning is proposed to analyze the behavior and decision-making of intelligent agents in complex environments. The designed network uses Q-learning to update the decision voltage in the circuit based on the optimal Bellman equation, allowing agents to make different strategies according to the constantly changing environment to achieve optimal results. In addition, the combination of Q-learning and memristors helps to alleviate the inherent overestimation bias in Q-learning, improve the stability and performance of the learning process. The proposed circuit can be applied in fields such as biomimetic robots, path planning and braking manufacturing, providing new ideas for the development of neuromorphic learning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Memristive Circuit Based on Q-Learning and Operant Conditioning

  • Lingying Kong,
  • Fei Fan,
  • Cunliang Zhang,
  • Qi’an Sun,
  • Junwei Sun,
  • Yanfeng Wang

摘要

In most memristive neural network circuits based on operant conditioning, the agent’s tendency towards certain behaviors is simply reflected through changes in synaptic weight. No specific analysis has been conducted on the changes in the tendency of intelligent agent behavior. Therefore, an operant conditioning circuit based on memristors and Q-learning is proposed to analyze the behavior and decision-making of intelligent agents in complex environments. The designed network uses Q-learning to update the decision voltage in the circuit based on the optimal Bellman equation, allowing agents to make different strategies according to the constantly changing environment to achieve optimal results. In addition, the combination of Q-learning and memristors helps to alleviate the inherent overestimation bias in Q-learning, improve the stability and performance of the learning process. The proposed circuit can be applied in fields such as biomimetic robots, path planning and braking manufacturing, providing new ideas for the development of neuromorphic learning.