Animal Learning in a Multidimensional Discrimination Task as Explained by Dimension-Specific Allocation of Attention

Animal Learning in a Multidimensional Discrimination Task as Explained by Dimension-Specific Allocation of Attention
复制标题

DOI:
10.3389/fnins.2018.00356
复制
发表时间:
2018-06-05
影响因子:
4.3
通讯作者:
Morris, Genela
Morris, Genela
中科院分区:
医学2区
文献类型:
--
作者:
Aluisi, Flavia;Rubinchik, Anna;Morris, Genela

文献摘要

被引文献

相似文献

强化学习描述了在一系列反复尝试的过程中,最终得到奖励的行动得到加强的过程,在类似的情况下变得更有可能被选择。当决策基于感官刺激时,刺激、行动和奖励之间就形成了一种联系。对这一过程的计算、行为和神经生物学解释成功地解释了对刺激的简单学习,这些刺激在一个方面或沿着单一刺激维度不同。然而,当刺激可能在几个维度上有所不同时,确定哪些特征与奖励相关并不是一件微不足道的事情,而且人们对潜在的认知过程也知之甚少。为了研究这一点,我们采用了维度内/维度外的集合转移范式来训练大鼠进行多感觉辨别任务。在我们的设置中,不同形式的刺激(空间、嗅觉和视觉)被组合成复杂的线索,并独立操作。在每一组中,只有一个刺激维度与奖励相关。为了区分学习和决策,我们建议使用加权注意力模型(WAM)。我们的模型通过为每个维度(例如,每个颜色)的特征值分配单独的学习规则来学习,并在每次体验后得到强化。决策是通过比较学习到的值的加权平均值来做出的,该加权平均值由特定维度的权重决定。根据观察到的大鼠行为,我们估计了WAM的参数,并证明它的表现优于另一种模型,在该模型中,每个特征组合都分配了一个学习值。WAM的估计决策权重揭示了学习中基于经验的偏见。在第一组实验中,与所有维度相关的权重都是相似的。超维度的移动使这个维度变得无关紧要。然而,在这最后一盘的早期学习阶段,它的决策权重仍然很高,这为动物表现不佳提供了一个解释。因此,估计的权重可以被视为量化基于经验的偏差的一种可能方法。
Reinforcement learning describes the process by which during a series of trial-and-error attempts, actions that culminate in reward are reinforced, becoming more likely to be chosen in similar circumstances. When decisions are based on sensory stimuli, an association is formed between the stimulus, the action and the reward. Computational, behavioral and neurobiological accounts of this process successfully explain simple learning of stimuli that differ in one aspect, or along a single stimulus dimension. However, when stimuli may vary across several dimensions, identifying which features are relevant for the reward is not trivial, and the underlying cognitive process is poorly understood. To study this we adapted an intra-dimensional/ extra-dimensional set-shifting paradigm to train rats on a multi-sensory discrimination task. In our setup, stimuli of different modalities (spatial, olfactory and visual) are combined into complex cues and manipulated independently. In each set, only a single stimulus dimension is relevant for reward. To distinguish between learning and decision-making we suggest a weighted attention model (WAM). Our model learns by assigning a separate learning rule for the values of features of each dimension (e.g., for each color), reinforced after every experience. Decisions are made by comparing weighted averages of the learnt values, factored by dimension specific weights. Based on the observed behavior of the rats we estimated the parameters of the WAM and demonstrated that it outperforms an alternative model, in which a learnt value is assigned to each combination of features. Estimated decision weights of the WAM reveal an experience-based bias in learning. In the first experimental set the weights associated with all dimensions were similar. The extra-dimensional shift rendered this dimension irrelevant. However, its decision weight remained high for the early learning stage in this last set, providing an explanation for the poor performance of the animals. Thus, estimated weights can be viewed as a possible way to quantify the experience-based bias.