Active inference and epistemic value

Active inference and epistemic value
复制标题

DOI:
10.1080/17588928.2015.1020053
复制
发表时间:
2015-10-02
影响因子:
2
通讯作者:
Pezzulo, Giovanni
Pezzulo, Giovanni
中科院分区:
医学4区
文献类型:
--
作者:
Friston, Karl;Rigoli, Francesco;Pezzulo, Giovanni

文献摘要

被引文献

相似文献

我们提供了一种选择行为的正式治疗,其前提是代理人将未来结果的预期自由能最小化。至关重要的是,政策的负自由能或质量可以分解为外在价值和认知(或内在)价值。因此,最小化预期自由能等同于最大化外在价值或预期效用(根据先前的偏好或目标定义),同时最大化信息收益或内在价值(或减少有价值结果原因的不确定性)。由此产生的方案解决了勘探-开发的两难问题:认知价值最大化,直到没有进一步的信息收益,之后通过最大化外部价值来确保开发。这在形式上与Infomax原理一致,概括了基于显著(贝叶斯惊喜)的主动视觉公式以及基于预期效用和风险敏感(Kullback-Leibler)控制的最佳决策。此外,与以前的离散(马尔可夫)问题的主动推理公式一样,特别Softmax参数成为关于策略的信念或信心的预期(贝叶斯最优)精度。本文重点阐述了该方法的基本理论,并通过仿真对其思想进行了说明。这些模拟的一个关键方面是在条件反射范例中观察到的精确更新和多巴胺能放电之间的相似性。
We offer a formal treatment of choice behavior based on the premise that agents minimize the expected free energy of future outcomes. Crucially, the negative free energy or quality of a policy can be decomposed into extrinsic and epistemic (or intrinsic) value. Minimizing expected free energy is therefore equivalent to maximizing extrinsic value or expected utility (defined in terms of prior preferences or goals), while maximizing information gain or intrinsic value (or reducing uncertainty about the causes of valuable outcomes). The resulting scheme resolves the exploration-exploitation dilemma: Epistemic value is maximized until there is no further information gain, after which exploitation is assured through maximization of extrinsic value. This is formally consistent with the Infomax principle, generalizing formulations of active vision based upon salience (Bayesian surprise) and optimal decisions based on expected utility and risk-sensitive (Kullback-Leibler) control. Furthermore, as with previous active inference formulations of discrete (Markovian) problems, ad hoc softmax parameters become the expected (Bayes-optimal) precision of beliefs about, or confidence in, policies. This article focuses on the basic theory, illustrating the ideas with simulations. A key aspect of these simulations is the similarity between precision updates and dopaminergic discharges observed in conditioning paradigms.