A comparison of reinforcement learning models of human spatial navigation.

A comparison of reinforcement learning models of human spatial navigation.
复制标题

DOI:
10.1038/s41598-022-18245-1
复制
发表时间:
2022-08-17
期刊:
影响因子:
4.6
通讯作者:
Brown, Thackery, I
Brown, Thackery, I
中科院分区:
综合性期刊3区
文献类型:
--
作者:
He, Qiliang;Liu, Jancy Ling;Eschapasse, Lou;Beveridge, Elizabeth H.;Brown, Thackery, I

文献摘要

参考文献

相似文献

强化学习(RL)模型在表征人类学习和决策方面具有重要影响,但很少有研究将其应用于表征人类空间导航,更少有研究系统地比较不同导航需求下的RL模型。由于强化学习可以定量地、连续地表征学习者的学习策略以及使用这些策略的一致性,因此它可以为理解人类导航中显著的个体差异和将导航策略与导航性能分开提供一个新的、重要的视角。114名参与者在虚拟环境中完成寻路任务,不同阶段操纵导航要求。我们比较了五种RL模型(3种无模型,1种基于模型和1种“混合”)在不同阶段拟合导航行为的性能。支持从以前的文献的影响,混合模型提供了最佳的适合,无论导航要求,这表明大多数参与者依赖于混合的无模型(路线跟踪)和基于模型(认知映射)的学习在这样的导航场景。此外,与关键预测一致,混合模型中基于模型的学习的权重(即,导航策略)和导航者的探索与利用倾向(即,使用这种导航策略的一致性),这是由导航任务要求调制。总之,我们不仅展示了RL的计算结果与空间导航文献的一致性,而且还揭示了导航策略和使用这些策略的人的一致性之间的关系如何随着导航需求的变化而变化。
Reinforcement learning (RL) models have been influential in characterizing human learning and decision making, but few studies apply them to characterizing human spatial navigation and even fewer systematically compare RL models under different navigation requirements. Because RL can characterize one’s learning strategies quantitatively and in a continuous manner, and one’s consistency of using such strategies, it can provide a novel and important perspective for understanding the marked individual differences in human navigation and disentangle navigation strategies from navigation performance. One-hundred and fourteen participants completed wayfinding tasks in a virtual environment where different phases manipulated navigation requirements. We compared performance of five RL models (3 model-free, 1 model-based and 1 “hybrid”) at fitting navigation behaviors in different phases. Supporting implications from prior literature, the hybrid model provided the best fit regardless of navigation requirements, suggesting the majority of participants rely on a blend of model-free (route-following) and model-based (cognitive mapping) learning in such navigation scenarios. Furthermore, consistent with a key prediction, there was a correlation in the hybrid model between the weight on model-based learning (i.e., navigation strategy) and the navigator’s exploration vs. exploitation tendency (i.e., consistency of using such navigation strategy), which was modulated by navigation task requirements. Together, we not only show how computational findings from RL align with the spatial navigation literature, but also reveal how the relationship between navigation strategy and a person’s consistency using such strategies changes as navigation requirements change.
DOI: 10.1146/annurev-psych-122414-033625
发表时间: 2017-01-03
影响因子: 24.8
作者:
Gershman SJ;Daw ND
通讯作者: Daw ND
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1016/s0160-2896(02)00116-2
发表时间: 2002-01-01
期刊: INTELLIGENCE
影响因子: 3
作者:
Hegarty, M;Richardson, AE;Subbiah, I
通讯作者: Subbiah, I
DOI: 10.1523/eneuro.0346-16.2017
发表时间: 2017-03
期刊: eNeuro
影响因子: 3.4
作者:
Chrastil ER;Sherrill KR;Aselcioglu I;Hasselmo ME;Stern CE
通讯作者: Stern CE
DOI: 10.3758/s13421-018-0811-y
发表时间: 2018-08-01
期刊: MEMORY & COGNITION
影响因子: 2.4
作者:
Boone, Alexander P.;Gong, Xinyi;Hegarty, Mary
通讯作者: Hegarty, Mary