Experience resetting in reinforcement learning facilitates exploration?exploitation transitions during a behavioral task for primates

Experience resetting in reinforcement learning facilitates exploration?exploitation transitions during a behavioral task for primates
复制标题

强化学习中的经验重置有助于灵长类动物行为任务中的探索和利用转变

DOI:
10.1101/2021.09.30.462676
复制
发表时间:
2021
期刊:
bioRxiv
影响因子:
--
通讯作者:
Mushiake Hajime
Mushiake Hajime
中科院分区:
--
文献类型:
--
作者:
Sakamoto Kazuhiro;Okuzaki Hidetake;Sato Akinori;Mushiake Hajime

文献摘要

相似文献

探索-利用权衡是再强化学习中的一个基本问题。为了研究涉及这个问题的神经机制,一个目标搜索任务,探索和开发阶段交替出现是有用的。在这个任务中受过良好训练的猴子清楚地知道,他们已经进入了探索阶段,并通过重置他们以前的经验迅速获得新的经验。在这项研究中,我们使用了一个简单的模型来表明,在探索阶段的经验重置提高性能,而不是减少贪婪的行动选择,然后我们提出了一个神经网络型模型,使经验重置。
The exploration–exploitation trade-off is a fundamental problem in re-inforcement learning. To study the neural mechanisms involved in this problem, a target search task in which exploration and exploitation phases appear alternately is useful. Monkeys well trained in this task clearly understand that they have entered the exploratory phase and quickly acquire new experiences by resetting their previous experiences. In this study, we used a simple model to show that experience resetting in the exploratory phase improves performance rather than decreasing the greediness of action selection, and we then present a neural network-type model enabling experience resetting.