Stimulus selection in a Q-learning model using Fisher information and Monte Carlo simulation

Stimulus selection in a Q-learning model using Fisher information and Monte Carlo simulation
复制标题

使用 Fisher 信息和蒙特卡罗模拟的 Q 学习模型中的刺激选择

DOI:
10.1007/s42113-022-00163-0
复制
发表时间:
2023
期刊:
Computational Brain & Behavior
影响因子:
--
通讯作者:
K.
K.
中科院分区:
--
文献类型:
--
作者:
Fujita;K;Okada;K;& Katahira;K.

文献摘要

相似文献

强化学习模型在具有奖励反馈的决策任务中得到了广泛的研究。然而,在设计Q-learning模型数据收集实验时,通常没有考虑所呈现的刺激对参与者参数估计精度的定量影响。也就是说,由于缺乏数学框架,研究人员无法设计出最佳的实验。为了解决这一问题,本研究对费雪信息进行了解析推导。此外,本研究还建立了Q-learning模型的随机表示,Q-learning模型是最常用的强化学习模型之一。在此基础上,提出了一种结合低成本Fisher信息评估和更详细的有限样本蒙特卡罗模拟的两步方法,以估计精度选择最优的刺激。仿真研究表明,奖励概率反转可以提高学习率参数的估计精度。相反,对于逆温度参数,选项之间奖励概率的差异越大,估计精度越高。这些结果表明,最优的实验设计取决于研究人员对q -学习模型的哪些特征参数感兴趣。此外,我们还发现,在性状参数精度方面使用不良刺激会导致相关系数估计的较大偏差。在此基础上,讨论了q -学习模型中实验设计的方法。
Reinforcement learning models have been extensively studied for decision-making tasks with reward feedback. However, in designing an experiment to collect data for Q-learning models, the quantitative effect of a presented stimulus on the estimation precision of participant parameters has generally not been considered. That is, the lack of a mathematical framework has prevented researchers from designing an optimal experiment. To tackle this problem, this study analytically derives the Fisher information. Furthermore, this study formulates a stochastic representation of the Q-learning model, which is one of the most commonly applied reinforcement learning models. With this derivation, a two-step procedure is proposed to select the optimal stimuli in terms of estimation precision, in which low-cost Fisher information evaluation and more detailed finite-sample Monte Carlo simulation are combined. The simulation studies show that reward probability reversal leads to a high estimation precision for the learning rate parameter. By contrast, for the inverse temperature parameter, a larger difference in reward probability between options leads to higher estimation precision. These results reveal that the optimal experimental design is dependent on which trait parameters of the Q-learning model are of interest to researchers. Further, it is found that the use of undesirable stimuli in terms of trait parameter precision leads to a large bias in the correlation coefficient estimate. Based on the results, the approaches to designing experiments in the Q-learning model are discussed.