课题基金 / 基金详情

EAGER: Income Learning: A New Model for Behavior-Analysis-Inspired Learning from Human Feedback

EAGER: Income Learning: A New Model for Behavior-Analysis-Inspired Learning from Human Feedback
EAGER:收入学习:基于人类反馈的行为分析启发学习的新模型
批准号:
1643614
负责人:
Matthew Taylor
金额:
$7.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-15 至 2017-07-31

项目摘要

项目成果

Matthew Taylor的其他基金

相似基金

相关文献

中文摘要
翻译
随着虚拟代理和物理机器人变得越来越普遍,它们可以有效地执行越来越多的复杂任务来帮助人类。这些任务通常被形式化为顺序决策任务,其中机器人和代理感知状态、采取行动并接收奖励反馈信号。在实践中,如果这样的机器要完成原始开发预先指定的任务之外的任务,就迫切需要直接从人类用户那里学习。机器强化学习(RL)是一种经常用于解决顺序决策任务的范式,最初是受应用行为分析(ABA)社区的动物学习研究的启发而发展起来的。现有的RL方法有效地操作了一组有限的ABA原则;然而,来自ABA研究的其他原则和性质没有很好地封装在现有的RL形式中,这些可能是设计能够从人类教师那里学习的更有效的RL技术的新灵感的来源。这个项目将(1)结合ABA和RL的原则来产生能够更有效地从人类那里学习的算法,(2)在虚拟代理和机器人平台上评估这些算法,以及(3)调查非专家人类是否以及如何构建难度越来越高的任务序列,类似于专家动物训练师如何塑造任务。来自这些用户研究的见解将被用来进一步提高我们的算法向人类训练者学习的能力。一旦成功,该项目将在允许非技术用户能够教授虚拟和物理代理人在自然环境中执行复杂任务方面取得关键进展,这是许多人从以前训练家养宠物的经验中熟悉的。该项目是华盛顿州立大学(WSU)、北卡罗来纳州立大学和布朗大学之间更大努力的一部分。华盛顿州立大学的工作将集中于实施拟议的机器学习算法家族,称为收入学习(i-Learning)。由于这些算法是由三所大学共同开发的,西斯安那州立大学将设计用户研究,以评估i-Learning背后的原则何时以及如何使其在从人类反馈中学习方面优于其他现有算法。WSU将主要关注1)虚拟代理,允许通过众包进行测试学习,以及在2)物理机器人上进行测试,并研究实施是否改变了用户的感知和行动,或算法的学习效率。此外,威斯康星州立大学将调查人类的课程设计。专家训练师可以塑造动物的行为,随着时间的推移,任务的复杂性会增加,因此动物学习一系列任务的速度比直接训练最后一项困难任务的速度要快得多。华盛顿州立大学将在众包平台上进行用户研究,以更好地了解非专家人类如何在连续决策任务中为机器学习算法设计课程,并调查这些设计决策如何影响算法设计。
英文摘要
As virtual agents and physical robots become more common, there is an increasing number of complex tasks they can usefully perform to assist humans. These tasks are typically formalized as sequential decision tasks, where robots and agents perceive states, take actions, and receive a reward feedback signal. In practice, there is a critical need to learn directly from human users if such machines are to accomplish tasks outside of those pre-specified by the original developments. Machine reinforcement learning (RL), a paradigm often used for solving sequential decision making tasks, was originally developed with inspiration from animal learning research from the applied behavior analysis (ABA) community. Existing RL approaches operationalize a limited set of ABA principles effectively; however, there are additional principles and properties from ABA research that are not well encapsulated in the existing RL formalisms, and that are likely sources of new inspiration for designing more effective RL techniques capable of learning from human teachers. This project will (1) take combine principles from ABA and RL to produce algorithms that can learn more effectively from humans, (2) evaluate these algorithms in both virtual agents and on robot platforms, and (3) investigate whether and how non-expert humans can construct sequences of tasks of increasing difficulty, similar to how expert animal trainers shape tasks. Insights from these user studies will be leveraged to further improve our algorithms' abilities to learn from human trainers. Once successful, this project will make critical progress towards allowing non-technical users to be able to teach virtual and physical agents to perform complex tasks in a natural setting, familiar to many from previous experience in training household pets.This project is a part of a larger effort between Washington State University (WSU), North Carolina State University, and Brown University. The WSU effort will focus on implementing the proposed family of machine learning algorithms, called Income Learning (I-Learning). As these algorithms are co-developed by the three universities, WSU will design user studies to evaluate when and how the principles behind I-Learning allow it to outperform other existing algorithms at learning from human feedback. WSU will primarily focus on 1) virtual agents, allowing test learning via crowdsourcing, as well as testing on 2) physical robots and study if embodiment changes user's perceptions and actions, or the algorithms' learning efficacy. Additionally, WSU will investigate 3) human curricula design. Expert trainers can shape the behavior of animals, increasing task complexity over time, so that the animals can learn a sequence of tasks much faster than if they trained directly on the final, difficult task. WSU will run user studies on crowdsourcing platforms to better understand how non-expert humans design curricula for machine learning algorithms in sequential decision tasks, and investigate how these design decisions can inform algorithm design.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DISES: Indigenous forest management in a non-stationary climate
  • 批准号:
    2310797
  • 项目类别:
    Standard Grant
  • 资助金额:
    $159.94万
  • 财政年份:
    2023
  • 负责人:
    Matthew Taylor
  • 依托单位:
Pilot study to develop a novel model to investigate the mechanisms and consequences of foetal immune programming on immune fitness through life
  • 批准号:
    BB/S002987/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $27.34万
  • 财政年份:
    2018
  • 负责人:
    Matthew Taylor
  • 依托单位:
Doctoral Mentoring Consortium at the Fourteenth International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-16)
  • 批准号:
    1620841
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.8万
  • 财政年份:
    2016
  • 负责人:
    Matthew Taylor
  • 依托单位:
19th Annual SIGART/AAAI Doctoral Consortium
海外基金