课题基金 / 基金详情

RI: Small: Coordination in tightly coupled domains: Stepping stone rewards to induce the correct joint actions

RI: Small: Coordination in tightly coupled domains: Stepping stone rewards to induce the correct joint actions
RI:小:紧密耦合领域中的协调:垫脚石奖励以诱导正确的联合行动
批准号:
1815886
负责人:
Kagan Tumer
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2023-08-31

项目摘要

项目成果

Kagan Tumer的其他基金

相似基金

相关文献

中文摘要
翻译
该项目引入了一种新的多智能体学习方法,该方法在紧密耦合的领域中导致协调行为,也就是说,在所有智能体必须在正确的时间做正确的事情以使团队实现其目标的领域中。例如,让一组代理抬起和移动一个比任何单个代理的有效载荷能力更重的对象,需要足够数量的代理在正确的时间执行正确的操作。不幸的是,大多数当前的学习方法在这种情况下都失败了,因为它们只依赖于在代理偶然发现正确的行为之后才加强正确的代理行为。但是,如果代理们不能共同找到正确的行动呢?这个项目通过引入“踏脚石奖励”来解决这个问题,即激励代理执行正确的行动,即使他们的队友还没有找到正确的补充行动。该项目的影响将是创建更大、更有能力的多智能体团队,这些团队可以部署在工业(如不限于单一任务的工厂机器人)、现场(如自主搜索和救援系统)、教育(如通过在线游戏进行互动学习)和家庭(如智能家电网络)中。这个项目的主要技术贡献是将智能体所面临的学习问题从“我是否采取了正确的行动?”转变为“如果其他智能体采取了补充行动,我的行动是否正确?”在紧密耦合的多智能体领域中,第一个问题产生的正反馈非常少,从而产生了一个很难甚至不可能的学习问题。新的垫脚石奖励利用假设的合作伙伴(由代理推测的合作伙伴,以探索联合行动空间),通过评估特定行动的潜在利益来克服这一困难。直观地说,垫脚石奖励为智能体创建了一个梯度,以便在紧密耦合的领域中实现快速有效的学习。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project introduces a new multiagent learning approach that leads to coordinated behavior in tightly coupled domains, that is, in domains where all agents must do the right thing at the right time for the team to achieve its goals. For example, getting a team of agents to lift and move an object heavier than the payload capacity of any single agent requires a sufficient number of agents to perform the correct action at the correct time. Unfortunately, most current learning methods fail in such situations because they rely on reinforcing the correct agent behavior only after the agents stumble upon the right actions. But what if the agents never jointly find the right actions? This project addresses this issue by introducing "stepping-stone rewards" that incentivize agents to perform the right actions even if their teammates have not yet found the correct complementary actions. The impact of this project will be to create larger and more capable multiagent teams that can be deployed in industry (such as factory robots that are not limited to a single task), in the field (such as autonomous search and rescue systems), in education (such as interactive learning via online gameplay) and in the home (such as networks of smart appliances).The main technical contribution of this project is to shift the learning problem faced by an agent from "did I take the correct action?" to "would my action have been correct had other agents taken the complementary action?" In tightly coupled multiagent domains, the first question results in very little positive feedback, creating a difficult to impossible learning problem. The new stepping stone rewards leverage hypothetical partners (partners that are surmised by an agent to explore the joint-action space) to overcome this difficulty by assessing the potential benefits of a particular action. Intuitively, stepping-stone rewards create a gradient for the agents to follow to enable fast and efficient learning in tightly coupled domains.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(16)
专著(0)
科研奖励(0)
会议论文
Bootstrapped fitness critics with bidirectional temporal difference
具有双向时间差异的自举健身批评家
DOI: 10.1145/3520304.3528999
发表时间: 2022
期刊: Genetic and Evolutionary Computation Conference
影响因子: --
作者: [Rockefeller, Golden, Tumer, Kagan]
通讯作者: Tumer, Kagan
Entropy-based local fitnesses for evolutionary multiagent systems
进化多智能体系统的基于熵的局部适应度
DOI: 10.1145/3520304.3529035
发表时间: 2022
期刊: Genetic and Evolutionary Computation Conference
影响因子: --
作者: [Aydeniz, Ayhan Alp, Nickelson, Anna, Tumer, Kagan]
通讯作者: Tumer, Kagan
Dynamic Skill Selection for Learning Joint Actions (extended abstract)
用于学习联合动作的动态技能选择(扩展摘要)
DOI: --
发表时间: 2021
期刊: Autonomous agents and multiagent systems
影响因子: --
作者: [Enna Sachdeva, Shauharda Khadka]
通讯作者: Enna Sachdeva, Shauharda Khadka
Diversifying behaviors for learning in asymmetric multiagent systems
非对称多智能体系统中学习行为的多样化
DOI: 10.1145/3512290.3528860
发表时间: 2022
期刊: Genetic and Evolutionary Computation Conference
影响因子: --
作者: [Dixit, Gaurav, Gonzalez, Everardo, Tumer, Kagan]
通讯作者: Tumer, Kagan
16
    Doctoral Mentoring Consortium at the Thirteenth International Conference on Autonomous Agents and Multi-Agent Systems
    • 批准号:
      1414600
    • 项目类别:
      Standard Grant
    • 资助金额:
      $2.5万
    • 财政年份:
      2014
    • 负责人:
      Kagan Tumer
    • 依托单位:
    CPS: Small: Collaborative Research: Distributed Coordination of Agents For Air Traffic Flow Management
    • 批准号:
      0931591
    • 项目类别:
      Standard Grant
    • 资助金额:
      $37.0万
    • 财政年份:
      2009
    • 负责人:
      Kagan Tumer
    • 依托单位:
    SGER: Foundations of Multiagent Control in Complex Environments
    • 批准号:
      0910358
    • 项目类别:
      Standard Grant
    • 资助金额:
      $12.88万
    • 财政年份:
      2009
    • 负责人:
      Kagan Tumer
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: