课题基金 / 基金详情

CPS: Medium: Sufficient Statistics for Learning Multi-Agent Interactions

CPS: Medium: Sufficient Statistics for Learning Multi-Agent Interactions
CPS:中:学习多智能体交互的足够统计数据
批准号:
2125511
负责人:
Dorsa Sadigh
金额:
$111.42万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-15 至 2025-08-31

项目摘要

项目成果

Dorsa Sadigh的其他基金

相似基金

相关文献

中文摘要
翻译
多智能体协调和协作是未来网络物理系统的核心挑战,因为它们开始彼此之间或与家庭或城市中的人类进行更复杂的交互。其中一个关键的挑战是,代理必须能够推理和学习其他代理的行为,以便能够做出决策。这尤其具有挑战性,因为最先进的方法,如合作伙伴策略上的递归信念建模,通常无法扩展。然而,人类在相互协调和协作方面非常有效,不需要任何昂贵的递归信念建模。一种假设是,人类可以有效地捕捉到协调任务所需的充分表征。与人类类似,多代理环境中的代理可以查找协调和协作所需的足够统计信息。这个项目是关于学习和接近这种充分的统计数据,以实现有效的协作和协调。此外,调查人员将在代理对世界进行部分观察的环境中研究教学和学习,需要相互教授和学习,以实现协作任务。单个代理的强化学习的重要演示成功,促使人们决定这种方法是否可以扩展到多个代理。在理解所产生的相互作用动力学的结构和开发实用的强化学习算法方面,多智能体系统领域也有了显著的发展。该项目的核心目标是:1)开发在多智能体相互作用中近似于众所周知的充分统计学概念的学习方法;2)开发一种强化学习算法,该算法利用充分统计学的表示法在多智能体环境中更有效地规划、协调和协作;以及3)开发使用充分统计学的表示法的算法,以便能够在环境中的部分观察下在多智能体环境中进行教和学。该项目的总体结果将是一种新的形式主义,以及增强多代理学习和控制的算法、工具和技术。调查人员将以两个主要应用为基础:1)合作搜索和探索以及2)合作运输物体。这一裁决反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Multi-agent coordination and collaboration is a core challenge of future cyber-physical systems as they start having more complex interactions with each other or with humans in homes or cities. One of the key challenges is that agents must be able to reason about and learn the behavior of other agents in order to be able to make decisions. This is particularly challenging because state of the art approaches such as recursive belief modeling over partner policies often do not scale. However, humans are very effective in coordinating and collaborating with each other without the need of any expensive recursive belief modeling. One hypothesis is that humans can effectively capture the sufficient representations required for coordinating on tasks. Similar to humans, the agents in a multi-agent setting can look for the sufficient statistics needed for coordination and collaboration. This project is about learning and approximating such sufficient statistics to enable effective collaboration and coordination. In addition, the investigators will study teaching and learning in settings where the agents have partial observation over the world and need to teach and learn from each other in order to achieve a collaborative task.Important successful demonstrations of reinforcement learning for single agents have spurred the drive to determine whether such methods can extend to multiple agents. There have also been notable developments in the area of multi-agent systems, both in understanding the structure of the resulting interacting dynamics and in the development of practical reinforcement learning algorithms. The core objective of this project is: 1) the development of learning methods that approximate the well-known concept of sufficient statistics in multi-agent interactions; 2) the development of a reinforcement learning algorithm that leverages the representations of sufficient statistics for more effective planning, coordination, and collaboration in multi-agent settings; and 3) the development of algorithms that use the representations of sufficient statistics to enable teaching and learning in multi-agent settings under partial observation over the environment. The overall outcome of this project will be a new formalism along with algorithms, tools, and techniques that enhance multi-agent learning and control. The investigators will ground this in two main applications: 1) collaborative search and exploration and 2) collaborative transport of objects.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/hri53351.2022.9889671
发表时间: 2022-01
期刊: 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI)
影响因子: --
作者: [Andy Shih;Stefano Ermon;Dorsa Sadigh]
通讯作者: Andy Shih;Stefano Ermon;Dorsa Sadigh
Partner-Aware Algorithms in Decentralized Cooperative Bandit Teams
去中心化合作强盗团队中的合作伙伴感知算法
DOI: --
发表时间: 2022
期刊: Proceedings of the 36th AAAI Conference on Artificial Intelligence
影响因子: --
作者: [Erdem Bıyık, Anusha Lalitha]
通讯作者: Erdem Bıyık, Anusha Lalitha
DOI: 10.48550/arxiv.2203.04421
发表时间: 2022-03
期刊: 2022 International Conference on Robotics and Automation (ICRA)
影响因子: --
作者: [Zhangjie Cao;Erdem Biyik;G. Rosman;Dorsa Sadigh]
通讯作者: Zhangjie Cao;Erdem Biyik;G. Rosman;Dorsa Sadigh
DOI: 10.48550/arxiv.2303.00001
发表时间: 2023-02
期刊: ArXiv
影响因子: --
作者: [Minae Kwon;Sang Michael Xie;Kalesha Bullard;Dorsa Sadigh]
通讯作者: Minae Kwon;Sang Michael Xie;Kalesha Bullard;Dorsa Sadigh
共 6 条
    Collaborative Research: CPS: Small: Risk-Aware Planning and Control for Safety-Critical Human-CPS
    • 批准号:
      2218760
    • 项目类别:
      Standard Grant
    • 资助金额:
      $25.0万
    • 财政年份:
      2022
    • 负责人:
      Dorsa Sadigh
    • 依托单位:
    NRI/Collaborative Research: Robot-Assisted Feeding: Towards Efficient, Safe, and Personalized Caregiving Robots
    • 批准号:
      2132847
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.85万
    • 财政年份:
      2022
    • 负责人:
      Dorsa Sadigh
    • 依托单位:
    Collaborative Research: Mixed-Autonomy Traffic Networks: Routing Games and Learning Human Choice Models
    • 批准号:
      1953032
    • 项目类别:
      Standard Grant
    • 资助金额:
      $18.0万
    • 财政年份:
      2020
    • 负责人:
      Dorsa Sadigh
    • 依托单位:
    CHS: Small: Learning and Leveraging Conventions in Human-Robot Interaction
    • 批准号:
      2006388
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2020
    • 负责人:
      Dorsa Sadigh
    • 依托单位:
    海外基金