<p>Choice behaviour of animals is characterized by two main tendencies: taking actions that led to rewards and repeating past actions<sup><CitationRef CitationID="CR1">1</CitationRef>,<CitationRef CitationID="CR2">2</CitationRef></sup>. Theory suggests that these strategies may be reinforced by different types of dopaminergic teaching signals: reward prediction error to reinforce value-based associations and movement-based action prediction errors to reinforce value-free repetitive associations<sup><CitationRef AdditionalCitationIDS="CR4 CR5" CitationID="CR3">3</CitationRef>–<CitationRef CitationID="CR6">6</CitationRef></sup>. Here we use an auditory discrimination task in mice to show that movement-related dopamine activity in the tail of the striatum encodes the hypothesized action prediction error signal. Causal manipulations reveal that this prediction error serves as a value-free teaching signal that supports learning by reinforcing repeated associations. Computational modelling and experiments demonstrate that action prediction errors alone cannot support reward-guided learning, but when paired with the reward prediction error circuitry they serve to consolidate stable sound–action associations in a value-free manner. Together we show that there are two types of dopaminergic prediction errors that work in tandem to support learning, each reinforcing different types of association in different striatal areas.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dopaminergic action prediction errors serve as a value-free teaching signal

  • Francesca Greenstreet,
  • Hernando Martinez Vergara,
  • Yvonne Johansson,
  • Sthitapranjya Pati,
  • Laura Schwarz,
  • Stephen C. Lenzi,
  • Jesse P. Geerts,
  • Matthew Wisdom,
  • Alina Gubanova,
  • Lars B. Rollik,
  • Jasvin Kaur,
  • Theodore Moskovitz,
  • Joseph Cohen,
  • Emmett Thompson,
  • Troy W. Margrie,
  • Claudia Clopath,
  • Marcus Stephenson-Jones

摘要

Choice behaviour of animals is characterized by two main tendencies: taking actions that led to rewards and repeating past actions1,2. Theory suggests that these strategies may be reinforced by different types of dopaminergic teaching signals: reward prediction error to reinforce value-based associations and movement-based action prediction errors to reinforce value-free repetitive associations36. Here we use an auditory discrimination task in mice to show that movement-related dopamine activity in the tail of the striatum encodes the hypothesized action prediction error signal. Causal manipulations reveal that this prediction error serves as a value-free teaching signal that supports learning by reinforcing repeated associations. Computational modelling and experiments demonstrate that action prediction errors alone cannot support reward-guided learning, but when paired with the reward prediction error circuitry they serve to consolidate stable sound–action associations in a value-free manner. Together we show that there are two types of dopaminergic prediction errors that work in tandem to support learning, each reinforcing different types of association in different striatal areas.