A Bayesian Account of Generalist and Specialist Formation Under the Active Inference Framework.

A Bayesian Account of Generalist and Specialist Formation Under the Active Inference Framework.
复制标题

主动推理框架下的通才和专家形成的贝叶斯帐户。

DOI:
10.3389/frai.2020.00069
复制
发表时间:
2020
影响因子:
4
通讯作者:
Friston KJ
Friston KJ
中科院分区:
其他
文献类型:
--
作者:
Chen AG;Benrimoh D;Parr T;Friston KJ

文献摘要

参考文献

被引文献

相似文献

本文在主动推理的框架下对策略学习或习惯行为优化进行了形式化的描述。在这种情况下,习惯的形成变成了一个自学的、依赖经验的过程,基于代理人看到自己在做什么。我们专注于环境波动对习惯形成的影响,通过模拟人工代理在一个部分可观察的马尔可夫决策过程。具体来说,我们使用了一个“两步”迷宫范式,其中代理人必须决定是向左还是向右走以获得奖励。我们观察到,在具有众多奖励位置的不稳定环境中,代理人学会采取通才策略,从未形成任何首选迷宫方向的强烈习惯行为。相反,在保守或静态的环境中,代理人采取专业的策略,形成强烈的政策偏好,导致接近少数先前观察到的奖励位置。两种策略的利弊进行了测试和讨论。一般来说,专业化提供了更大的好处,但只有当意外事件随着时间的推移而保存。我们认为这种正式的(主动推理)帐户的政策学习理解专业化和习惯形成之间的关系的影响。
This paper offers a formal account of policy learning, or habitual behavioral optimization, under the framework of Active Inference. In this setting, habit formation becomes an autodidactic, experience-dependent process, based upon what the agent sees itself doing. We focus on the effect of environmental volatility on habit formation by simulating artificial agents operating in a partially observable Markov decision process. Specifically, we used a “two-step” maze paradigm, in which the agent has to decide whether to go left or right to secure a reward. We observe that in volatile environments with numerous reward locations, the agents learn to adopt a generalist strategy, never forming a strong habitual behavior for any preferred maze direction. Conversely, in conservative or static environments, agents adopt a specialist strategy; forming strong preferences for policies that result in approach to a small number of previously-observed reward locations. The pros and cons of the two strategies are tested and discussed. In general, specialization offers greater benefits, but only when contingencies are conserved over time. We consider the implications of this formal (Active Inference) account of policy learning for understanding the relationship between specialization and habit formation.
DOI: 10.3389/fncom.2012.00024
发表时间: 2012
影响因子: 3.2
作者:
Fernando C;Szathmáry E;Husbands P
通讯作者: Husbands P
DOI: 10.1111/j.1553-2712.2008.00227.x
发表时间: 2008-11-01
影响因子: 4.4
作者:
Ericsson, K. Anders
通讯作者: Ericsson, K. Anders
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1162/neco_a_00912
发表时间: 2017-01-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Friston, Karl;FitzGerald, Thomas;Pezzulo, Giovanni
通讯作者: Pezzulo, Giovanni
DOI: 10.1016/j.neuroimage.2011.03.062
发表时间: 2011-06-15
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Friston, Karl J.;Penny, Will
通讯作者: Penny, Will