课题基金 / 基金详情

SHAIC1: Towards Scalable Human-AI Coordination from First Principles

SHAIC1: Towards Scalable Human-AI Coordination from First Principles
SHAIC1:从第一原则迈向可扩展的人类与人工智能协调
批准号:
EP/Y028481/1
负责人:
Jakob Foerster
金额:
$242.77万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这项提议的目标是开发人工智能(AI)代理,可以在复杂的现实世界环境中支持人类并与其协作。例如,这些机器人包括可以与人类合作的工业或服务机器人,以及在混合自动驾驶环境中与其他交通参与者顺利互动的自动驾驶汽车。一个根本的问题是,与竞争环境下的可扩展解决方案不同,目前的合作解决方案依赖于人类数据,因此在可扩展性方面受到限制。不幸的是,在这些环境中,扩展计算以消除对人工数据的需求是具有挑战性的。如果没有零和环境中明确定义的目标,就需要在一个池中找到为数不多的几个人类相容的解决方案之一,而这个池中还包含许多人类不相容的解决方案。我的假设是,人类有一个明确的概念,即良好的协调解决方案,并且在很大程度上依赖这个概念来解决协调问题,即当他们必须与他人合作,但无法预先就策略达成一致时。一般而言,此类场景中的好解决方案是简单、对称的解决方案,因此易于适应。为了将这个直观的想法形式化和实现,我将展示如何使用迭代学习的状态抽象来实现简单性和对称性约束,从而在复杂的环境中高效地发现通用协调策略。然后,我将使用通过在线适应或少量真实世界人类数据逐渐放松约束的新算法,将这些策略粗壮地扩展到人类次最优。这个项目将产生新的方法,可以扩展到当前技术水平所无法企及的复杂的人类-人工智能协调问题。它还将开发一种新的理论,为人类-人工智能协调方面的根本进展奠定基础,并开启关键应用领域,如自主工业机器人。
英文摘要
The goal of this proposal is to develop artificial intelligence (AI) agents, that can support and collaborate with humans in complex, real-world settings. These include, for example, industrial or service robots that can work in teams with humans and self-driving cars that interact smoothly with other traffic participants in mixed-autonomy settings. A fundamental issue is that, unlike the scalable solutions for competitive settings, current approaches for cooperative ones rely on human data and are thus limited in their scalability. Unfortunately, scaling compute to remove the need for human data is challenging in these settings. Without the well-defined objective present in zero-sum settings, it requires finding one of the few solutions that is human-compatible in a pool that also contains combinatorially many human-incompatible ones.My hypothesis is that humans have a well-defined concept of a 'good coordination solution' and to a great extent rely on this concept to solve coordination problems, i.e. when they have to work with others but cannot pre-agree on a strategy. Generally speaking, a good solution in such scenarios is one that is simple, symmetric, and therefore easy to adapt to. To move towards a formalisation and implementation of this intuitive idea, I will show how general purpose coordination policies can be efficiently discovered in complex settings using iteratively learned state-abstractions which implement simplicity and symmetry constraints.I will then robustify these policies to human sub-optimality using novel algorithms that gradually relax the constraints via online adaptation or small amounts of real-world human data.This project will result in new methods that can scale to complex human-AI coordination problems beyond the reach of the current state of the art. It will also develop a new theory that sets the scene for fundamental progress on human-AI coordination and unlocks crucial application areas, such autonomous industrial robots.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金