Using the Crowd to Prevent Harmful AI Behavior

Using the Crowd to Prevent Harmful AI Behavior
复制标题

DOI:
10.1145/3415168
复制
发表时间:
2020-10
影响因子:
--
通讯作者:
Travis Mandel;Jahnu Best;Randall H. Tanaka;Hiram Temple;Chansen Haili;Sebastian J. Carter;Kayla Schlechtinger;Roy Szeto
Travis Mandel;Jahnu Best;Randall H. Tanaka;Hiram Temple;Chansen Haili;Sebastian J. Carter;Kayla Schlechtinger;Roy Szeto
中科院分区:
--
文献类型:
--
作者:
Travis Mandel;Jahnu Best;Randall H. Tanaka;Hiram Temple;Chansen Haili;Sebastian J. Carter;Kayla Schlechtinger;Roy Szeto

文献摘要

相似文献

为了防止有害的人工智能行为,人们需要指定禁止不良行为的约束。不幸的是,这是一项复杂的任务,因为在现实世界中,编写区分有害和无害行为的规则往往相当困难。因此,这些决定历来都是由一小群强大的人工智能公司和开发者做出的,社区的投入有限。在本文中,我们研究了如何使一群非人工智能专家一起工作,向人工智能系统传达高质量、可靠的约束。我们首先专注于理解人类如何在人工智能行为的背景下对时间动态进行推理,通过在一个基于游戏的新型测试平台上的实验发现,参与者倾向于采用长期伤害的概念,即使是在不确定的情况下,也不会直接影响他们。在此基础上,我们探索了长期约束规范的任务设计,开发了新的过滤方法和促进用户反思的新方法。接下来,我们开发了一个新颖的基于规则的接口,它允许人们在没有编程知识的情况下以可访问的方式制作规则。我们在教育领域的一个真实的人工智能问题上测试了我们的方法,发现我们的新过滤机制和接口显着提高了约束质量和人类效率。我们还演示了这些系统如何应用于其他现实世界的人工智能问题(例如,在社交网络中)。
To prevent harmful AI behavior, people need to specify constraints that forbid undesirable actions. Unfortunately, this is a complex task, since writing rules that distinguish harmful from non-harmful actions tends to be quite difficult in real-world situations. Therefore, such decisions have historically been made by a small group of powerful AI companies and developers, with limited community input. In this paper, we study how to enable a crowd of non-AI experts to work together to communicate high-quality, reliable constraints to AI systems. We first focus on understanding how humans reason about temporal dynamics in the context of AI behavior, finding through experiments on a novel game-based testbed that participants tend to adopt a long-term notion of harm, even in uncertain situations that do not affect them directly. Building off of this insight, we explore task design for long-term constraint specification, developing new filtering approaches and new methods of promoting user reflection. Next, we develop a novel rule-based interface which allows people to craft rules in an accessible fashion without programming knowledge. We test our approaches on a real-world AI problem in the domain of education, and find that our new filtering mechanisms and interfaces significantly improve constraint quality and human efficiency. We also demonstrate how these systems can be applied to other real-world AI problems (e.g. in social networks).