EAGER: On the Optimal Rewards Problem
EAGER: On the Optimal Rewards Problem
批准号:
1148668
负责人:
Satinder Baveja
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-01 至 2014-07-31
中文摘要
以效用或奖励函数的形式指定目标是开发自主人工智能体以及理解和塑造自然生物智能体行为的大多数方法的基石--在控制、人工智能、经济学、心理学和行为学等领域。但在实践中,目标实际上有两个概念:(人类或进化的)设计者的目标和(人工或自然的)代理人的目标。这些应该是一样的吗?传统的(隐式)答案是“是”,但PI的新工作表明,对于计算有界的代理,答案可能是“否”。我们在智能体设计中定义了一个新的问题--最优报酬问题,它的解是分配给智能体的报酬函数,以使智能体在试图最大化其报酬时最大限度地实现设计者的目标。这个项目通过计算实验和分析,在两个广泛的方面探索最佳回报。我们将为有界计划代理开发新的原则和算法,展示出比使用传统奖励更好的性能,即使考虑到寻找最优奖励的成本。我们将在行为经济学和行为学的两个领域取得重大进展:理解人类的主观效用,以及理解动物的觅食行为,使用最优奖励理论严格推导考虑自然代理人计算限制的主观奖励函数。通过已发表的理论和软件传播的这项工作的结果,应该会导致我们理解和设计人类、动物和人工代理人的激励结构的方式发生根本性的变化,从而为工程、教育、经济和健康中依赖于为期望的行为结果找到良好激励的任何方法带来显著的实际好处。
英文摘要
The specification of goals in the form of utility or reward functions is a cornerstone of most approaches to developing autonomous artificial agents, and to understanding and shaping the behavior of natural biological agents - in fields ranging from control, AI, economics, psychology, and ethology. But in practice there are actually two notions of goals: the (human or evolutionary) designer's goals and the (artificial or natural) agent's goals. Should these be the same? The conventional (implicit) answer is "yes", but new work by the PIs shows that the answer may be "no" for computationally bounded agents. We define a new problem in agent design, the optimal rewards problem, whose solution is a reward function to assign to the agent so that in attempting to maximize its reward the agent best achieves the designer's goals. This project explores optimal rewards, through computational experimentation and analysis, on two broad fronts. We will develop new principles and algorithms for bounded planning agents, demonstrating increased performance over using conventional rewards, even when taking into account the cost of finding the optimal rewards. We will make significant advances in two areas of behavioral economics and ethology: the understanding of subjective utility in humans, and the understanding of foraging behavior in animals, by using optimal reward theory to rigorously derive subjective reward functions that take into account the computational limits of the natural agents. The results of this work, disseminated through published theory and software, should lead to foundational changes in the way we understand and design incentive structures for humans, animals, and artificial agents, and thus to significant practical benefits for any methods in engineering, education, economics, and health that depend on finding good incentives for desired behavioral outcomes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Combining Reinforcement Learning and Deep Learning Methods to Address High-Dimensional Perception, Partial Observability and Delayed Reward
-
批准号:1526059
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2015
-
负责人:Satinder Baveja
-
依托单位:
RI: Small: Reinforcement Learning with Predictive State Representations
-
批准号:1319365
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2013
-
负责人:Satinder Baveja
-
依托单位:
SHB: Medium: Collaborative Research: Novel Computational Techniques for Cardiovascular Risk Stratification
-
批准号:1064948
-
项目类别:Standard Grant
-
资助金额:$56.24万
-
财政年份:2011
-
负责人:Satinder Baveja
-
依托单位:
RI: Medium: Building Flexible, Robust, and Autonomous Agents
-
批准号:0905146
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Satinder Baveja
-
依托单位:
Flexible State Representations in Reinforcement Learning
-
批准号:0413004
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Satinder Baveja
-
依托单位:
Collaborative Research: Intrinsically Motivated Learning in Artificial Agents
-
批准号:0432027
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:2004
-
负责人:Satinder Baveja
-
依托单位:
Exploiting Structure in Reinforcement Learning Problems
-
批准号:9711753
-
项目类别:Continuing Grant
-
资助金额:$22.97万
-
财政年份:1997
-
负责人:Satinder Baveja
-
依托单位:
海外基金