<p>To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive behaviour can be learned through reinforcement learning<sup><CitationRef CitationID="CR1">1</CitationRef></sup>, a class of algorithms that has been successful at training artificial agents<sup><CitationRef AdditionalCitationIDS="CR3 CR4" CitationID="CR2">2</CitationRef>–<CitationRef CitationID="CR5">5</CitationRef></sup> and at characterizing the firing of dopaminergic neurons in the midbrain<sup><CitationRef AdditionalCitationIDS="CR7" CitationID="CR6">6</CitationRef>–<CitationRef CitationID="CR8">8</CitationRef></sup>. In classical reinforcement learning, agents discount future rewards exponentially according to a single timescale, known as the discount factor. Here we explore the presence of multiple timescales in biological reinforcement learning. We first show that reinforcement agents learning at a multitude of timescales possess distinct computational benefits. Next, we report that dopaminergic neurons in mice performing two behavioural tasks encode reward prediction error with a diversity of discount time constants. Our model explains the heterogeneity of temporal discounting in both cue-evoked transient responses and slower timescale fluctuations known as dopamine ramps. Crucially, the measured discount factor of individual neurons is correlated across the two tasks, suggesting that it is a cell-specific property. Together, our results provide a new paradigm for understanding functional heterogeneity in dopaminergic neurons and a mechanistic basis for the empirical observation that humans and animals use non-exponential discounts in many situations<sup><CitationRef AdditionalCitationIDS="CR10 CR11" CitationID="CR9">9</CitationRef>–<CitationRef CitationID="CR12">12</CitationRef></sup>, and open new avenues for the design of more-efficient reinforcement learning algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-timescale reinforcement learning in the brain

  • Paul Masset,
  • Pablo Tano,
  • HyungGoo R. Kim,
  • Athar N. Malik,
  • Alexandre Pouget,
  • Naoshige Uchida

摘要

To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive behaviour can be learned through reinforcement learning1, a class of algorithms that has been successful at training artificial agents25 and at characterizing the firing of dopaminergic neurons in the midbrain68. In classical reinforcement learning, agents discount future rewards exponentially according to a single timescale, known as the discount factor. Here we explore the presence of multiple timescales in biological reinforcement learning. We first show that reinforcement agents learning at a multitude of timescales possess distinct computational benefits. Next, we report that dopaminergic neurons in mice performing two behavioural tasks encode reward prediction error with a diversity of discount time constants. Our model explains the heterogeneity of temporal discounting in both cue-evoked transient responses and slower timescale fluctuations known as dopamine ramps. Crucially, the measured discount factor of individual neurons is correlated across the two tasks, suggesting that it is a cell-specific property. Together, our results provide a new paradigm for understanding functional heterogeneity in dopaminergic neurons and a mechanistic basis for the empirical observation that humans and animals use non-exponential discounts in many situations912, and open new avenues for the design of more-efficient reinforcement learning algorithms.