I'd do anything for a cookie (but I won't do that): Children's understanding of the costs and rewards underlying rational action

I'd do anything for a cookie (but I won't do that): Children's understanding of the costs and rewards underlying rational action
复制标题

为了一块饼干我愿意做任何事(但我不会那样做):孩子们对理性行为背后的成本和回报的理解

DOI:
--
复制
发表时间:
2014
期刊:
Annual Meeting of the Cognitive Science Society
影响因子:
--
通讯作者:
L. Schulz
L. Schulz
中科院分区:
--
文献类型:
--
作者:
J. Jara;H. Gweon;J. Tenenbaum;L. Schulz

文献摘要

被引文献

相似文献

我会为饼干做任何事情(但我不会那样做): *孩子们对成本和奖励的理解理性行动Julian Jara-Ettinger(jjara@mit.edu),Hyowon Gweon(hyora@mit.edu) ,Joshua B. Tenenbaum(jbt@mit.edu)和Laura E. Schulz(lschulz@mit.edu)大脑和认知科学系马萨诸塞州剑桥,马萨诸塞州剑桥,马萨诸塞州02139美国摘要Tenenbaum,2010年; Jara-Ettinger,Baker&Tenenbaum,2012年)。鉴于可能的世界状态的规范,代理商的可能行动及其可能的结果,以及代理商的实用程序功能2(正面和负面评价),最有效的方式是奖励与这些概率通用模型的不同组合和世界各州的不同组合可以实施一种理性的逆计划,从观察到代理商的行动来推断代理商的世界模型或公用事业函数规划帐户已被用来对成年人对代理人的欲望,信仰和世界的法官进行细粒度的定量预测(Baker等,2009,Baker,Saxe和Tenenbaum 2011; Jara-Ettinger等人,2012年)。 ,除了他们遗漏的内容,因为它们激励我们目前的工作。要实现的目标,衡量目标对代理商的价值,以及与可以实现这些目标的行动相关的(负)成本术语,以正式衡量行动的难度。通常,代理的状态和行动的联合功能)与每个状态相关的奖励,以及与每个动作相关的成本:u(a,s)= r(s)-c(a)。一个动作,a,为了实现状态,S仅表示S的相对奖励明显高于A的成本;高奖励,或者行动相对较低,目标状态在奖励中相当较低,可能同样可行的解释相同的行为。成本/低奖励。这样的推论是通过理性行动的原则来实现的:期望代理在情境约束中有效地实现其目标。动作是由幼稚的实用程序对代理恒定的敏感和特定于代理的成本方面以及与动作相关的奖励的敏感。奖励(即代理人的喜好),反之亦然,孩子们可以对物体和代理人设计信息性干预措施关于代理人的行为:幼稚的效用算术;代理人将采取最短的途径,即受世界施加的物理约束的目标(Gergely&Csibra,2003年)。强大的支持对未来事件的预测和关于事件的未观察到的推论,如果萨利在墙上跳过饼干,我们假设她不会跳跃,但是如果没有墙,那里的研究表明,即使是婴儿,婴儿也可以使用有关代理人的目标和情境约束的信息(例如,差距,遮挡者,墙壁等) Nadasdy,Csibra和biro,1995年); 2003年,请参见Brandone&Wellman,2009年; Carpenter,&Tomasello,2006年; Scott&Baillargeon,2013年。 Tenenbaum,2009年,2011年;但是,由于此功能是从奖励成本中得出的,因此我们将其称为“公用事业”功能。
I’d do anything for a cookie (but I won’t do that): * Children’s understanding of the costs and rewards underlying rational action Julian Jara-Ettinger (jjara@mit.edu), Hyowon Gweon (hyora@mit.edu), Joshua B. Tenenbaum (jbt@mit.edu), & Laura E. Schulz (lschulz@mit.edu) Department of Brain and Cognitive Sciences Massachusetts Institute of Technology, Cambridge, MA 02139 USA Abstract Tenenbaum, 2010; Jara-Ettinger, Baker, & Tenenbaum, 2012). MDPs are a framework widely used in artificial intelligence and other engineering fields for determining sequences of actions, or plans, an agent can take to achieve the highest-utility states in the most efficient manner, given a specification of the possible world states, the agent’s possible actions and their likely outcomes, and the agent’s utility function 2 (positively and negatively valued rewards) associated with different combinations of actions and world states. Bayesian inference over these probabilistic generative models can implement a form of rational inverse planning, working backwards from observations of an agent’s actions to infer aspects of the agent’s world model or utility function. Bayesian inverse planning accounts have been used to make fine-grained quantitative predictions of adults’ judgments about an agent’s desires, beliefs, and states of the world (Baker, et al., 2009, Baker, Saxe, & Tenenbaum 2011; Jara-Ettinger et al., 2012). The details of this computational approach are not critical here, but it is helpful to consider the qualitative intuitions behind these models, as well as what they leave out, because they motivate our present work. Intuitively we can think of an agent’s utility function as the difference between two terms: a (positive) reward term associated with goals to be achieved, measuring the value of a goal to the agent, and a (negative) cost term associated with actions that can be taken to achieve these goals, measuring the difficulty of an action. Formally, we can decompose the utility function (normally a joint function of the agent’s state and actions) into a reward associated with each state, and a cost associated with each action: U(a,s)=R(s)-C(a). Note that observing an agent taking an action, a, to achieve state, s, implies only that the relative reward for s is significantly higher than the cost of a; it does not determine either of these values in absolute terms: positing that the action has high cost but the goal state generates very high rewards, or that the action is relatively lower cost and the goal state is comparably lower in reward, maybe equally viable explanations of the same behavior. Psychologically however, high cost/high reward plans are very different from low cost/low reward ones. If Sally jumps over a wall to get a cookie is it because she likes the cookies so much Humans explain and predict other agents’ behavior using mental state concepts, such as beliefs and desires. Computational and developmental evidence suggest that such inferences are enabled by a principle of rational action: the expectation that agents act efficiently, within situational constraints, to achieve their goals. Here we propose that the expectation of rational action is instantiated by a naive utility calculus sensitive to both agent-constant and agent-specific aspects of costs and rewards associated with actions. We show that children can infer unobservable aspects of costs (differences in agents’ competence) from information about subjective differences in rewards (i.e., agents’ preferences) and vice versa. Moreover, children can design informative interventions on both objects and agents to infer unobservable constraints on agents’ actions. 1 Keywords: Naive Utility Calculus; Social Cognition; Theory of Mind Introduction One of the assumptions underlying our ability to draw rich inferences from sparse data is that agents act rationally. In its simplest form, this amounts to the expectation that agents will take the shortest path to a goal subject to physical constraints imposed by the world (Gergely & Csibra, 2003). Even this simple formulation is inferentially powerful, supporting predictions about future events and inferences about unobserved aspects of events. For instance, if Sally hops over a wall to get a cookie, we assume that she would not hop, but walk straight to the cookie, if the wall weren’t there. Studies suggest that even infants expect agents to act rationally. Infants can use information about an agent’s goal and situational constraints (e.g., gaps, occluders, walls, etc.) to predict her actions (Gergely, Nadasdy, Csibra, & Biro, 1995); an agent’s actions and situational constraints to infer her goals (Csibra, Biro, Koos, & Gergeley, 2003), and an agent’s actions and goals to infer unobserved situational constraints (see Csibra et al., 2003 for review; see also Brandone & Wellman, 2009; Gergeley, Bekkering, & Kiraly, 2002; Phillips & Wellman, 2005; Schwier, Van Maanen, Carpenter, & Tomasello, 2006; Scott & Baillargeon, 2013). Computationally, this approach to action understanding can be formalized as Bayesian inference over a model of rational action planning, such as a Markov Decision Process (MDP) (Baker, Saxe, & Tenenbaum, 2009, 2011; Ullman, Baker, Macindoe, Evans, Goodman, & In the artificial intelligence literature this is sometimes referred to as the reward function. However, since this function is derived from rewards minus costs, we refer to it as the utility function for clarity. * Or That’s the way the utility crumbles.